28edf7c75a made stage() chmod the native agent-browser binary, but
entries staged before it (lazy installs since the pm store landed) stay
0644 forever: every browser_* call fails with PermissionError and the
store never restages because the entry verifies. Restore the bit in
place when the entry is looked up on its own target (modes are not part
of the pinned digest); an entry that cannot be chmodded counts as not
installed so ensure() restages it.
The isolation filter now drops only settings that redirect WHERE packages
resolve from (index URLs, find-links, strategy, index credentials); transport
knobs (UV_NATIVE_TLS / UV_INSECURE_HOST / UV_HTTP_TIMEOUT) survive — corporate
networks need them to reach the pinned wheel URLs at all.
Addresses the review note on the no_config scope in the #125071 triage.
stage_runtime (the pinned pm/ self-repair builder) passes no_config=True,
but that only appended uv's --no-config, which removes config FILES. The
pip-mirror bridge (pm.index_config) reaches uv as env vars instead, so a
user's pip.conf index-url still rode into the runtime sync via
runtime_environment() -> _base_environment(env=None). Re-resolving the
official-index pm/uv.lock against a mirror changes the resolution, and
uv sync --locked refuses — breaking the top-level update / doctor /
pm-repair recovery paths on any machine with a pip mirror (#124418).
When no_config is set, strip forwarded index/transport knobs from the
child env: a fully-pinned graph never needs a mirror to resolve, and the
default path keeps the bridge so mirrored/air-gapped application installs
still work.
A time-based grace kept a just-abandoned lease 'held', so back-to-back `hermes pm repair` runs
never collected the generation the previous run execv'd out of (e2e test_generation_gc red).
Rename+replace is out (msvcrt byte locks refuse renames on Windows), so the writer now re-checks
its lease path after taking the flock and retakes when a peer's prune unlinked it in between; the
pruner only unlinks while holding the lock, so a visible path after the lock is always ours.
Review majors on #126179 (runtime_state._prune_unlocked_leases):
1. The boot-time prune opened every sibling lease with O_RDWR and caught
only FileNotFoundError, so one lease this user cannot open (a root-owned
0o600 file from `sudo hermes` on the same checkout, the #125525 class)
raised PermissionError inside activate_dependencies and hermes_bootstrap
exited 1 with "run hermes pm repair" on every launch. Any other OSError
now counts the lease as HELD: GC stays fail-closed, boot never aborts.
2. lease_directory creates (O_CREAT|O_EXCL) and flocks in two syscalls, and
pm/worker.py, pm/launch.py and pm/environments.py call it outside
runtime_lock, so a peer's prune in that window could take the
non-blocking lock and unlink the file; the first process then held a
lock on an unlinked inode and its pin was invisible for its lifetime
(collect_generations could rmtree a running generation). A lease younger
than LEASE_GRACE_SECONDS (60 s) is never a prune candidate and counts as
held. Chosen over a temp-name + os.replace scheme because it is three
lines, keeps a crash-between-create-and-flock leak collectable, and a
60 s margin dwarfs a microsecond window.
test_next_reader_removes_lease_left_by_hard_exit ages the hard-exited
lease past the grace window before asserting the next reader removes it;
its invariant is unchanged.
The runtime venv installs the core project editable against the PM
build snapshot (workspace/), so venv/bin/hermes always imports that
frozen copy. PM only rebuilds the venv on lock/extras/python-pin/plugin
changes, so a source-only update leaves the copy stale while the
gateway keeps running the checkout via activate_dependencies's
sys.path override. Any child that resolves `hermes` off PATH (terminal
tool, cron, kanban workers) still finds venv/bin first and runs the
stale snapshot instead — including `gateway install`/`restart`, which
then pins launchd/systemd to the snapshot's launcher path.
Prepend the checkout's own .hermes/bin ahead of venv/bin in
activate_dependencies so `hermes`/`hermes-acp` always resolve to the
launcher that puts the checkout first, regardless of what else lives
in the venv. This also fixes tools/environments/local.py's
_resolve_hermes_bin_dir(), whose shutil.which("hermes") now finds the
launcher dir instead of the venv's console script, so no separate guard
on PROJECT_ROOT in generate_launchd_plist / systemd unit generation is
needed: the snapshot's hermes_cli is never the running one.
Fixes#124627
lease_directory releases its per-invocation lease file only through atexit,
which os.execv (hermes_bootstrap's re-exec into the managed environment) and
os._exit (gateway shutdown_watchdog._hard_exit, hermes_cli/main.py past
finalization) never run: a 4s supervisor restart loop left ~860 lease files an
hour in the SELECTED generation, which generation GC never visits.
The kernel lock dies with the process even when atexit does not, so the file
itself is the only leak: every reader (the next lease_directory and
leases_held) now unlinks lease files nobody holds a lock on. Both run under
runtime_lock at boot / in the collector, so a lease between O_CREAT|O_EXCL and
its lock is never pruned from under a live process.
pm/lock.json and the Npm/AgentBrowser templates record registry.npmjs.org, and
pinned_source handed that URL straight to the downloader, so ~/.npmrc /
npm_config_registry were ignored and closed networks could not provision npm or
agent-browser. pinned_source now routes public-npm URLs through the configured
registry (npm's own precedence: npm_config_registry, then the user npmrc); the
lock's SHA256 still verifies the bytes and progress keeps the lockfile URL.
npm_dist_tags reads from the same registry.
The E2E cell test_npm_registry_mirror_serves_pm_npm_download drops its gate.
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
Post-write .js/.ts lint on the local backend runs PM's node/npx instead of
the terminal shell's PATH (user dirs first). pm.activate always puts the
store dirs in front, also on drift, activating only packages whose whole
chain is installed. tui_gateway puts the store dirs ahead of ~/.local/bin.
Bare npx/npm/node/uv/uvx MCP commands resolve to PM's copies; the system
fallback tables are gone and an absolute command: stays the user's.
Maintainer ruling: only Hermes and only its packaged package managers
(uv/pip/node/npm) are ever used; no PATH fallback when the managed tool is
missing, no "prefer the user's if new enough".
- hermes_constants.find_node_executable: node/npm/npx resolve to PM's
installed copy or None. Every caller already pm.ensure()s on None, so a
missing runtime is now provisioned instead of silently borrowing the
user's Node (native-addon ABI / npm cache mismatches).
- agent/lsp/install._install_npm: pm.ensure('npm') when PM npm is absent,
instead of failing over to whatever npm is on PATH.
- gateway._append_node_dir_for_service: stop baking the invoker's PATH node
dir into generated systemd/launchd units.
- main_install_repair._resolve_node_runtime_npm: drop the PATH re-scan for
another npm.
- source_build.source_product_current: run the freshness reader only with
PM's node.
- doctor: Node/npm rows and npm audit use PM's copies (Termux APT distro
keeps its system Node).
- install.sh ensure_uv / install.ps1 Get-Uv: always stage the pinned uv
artifact; delete the "uv on PATH if new enough" developer shortcut.
Computer use and browser use are meant to work out of the box. The PM rewrite
(3d12e86ef1) and the MSIX installer rework (47f4ab3a17) dropped the
install-time cua-driver fetch (7060ac7bed) and the Browser Use CLI install
(baa6b2e34d). On a fresh install the computer_use check_fn therefore stayed
False, so the tool never reached the model and its lazy ensure could not fire,
and browser_exec quietly fell back to the built-in tools.
- cua-driver is a default PM package, so the installers, a bare
`hermes pm install` and `hermes update` carry it on every target it builds
for. Adds the missing Android gap (the lock has no bionic artifact).
- The shared default-tool step, which the installers (via source completion)
and `hermes update` both run, provisions the Browser Use CLI for the default
and explicit Browser Use backends. `--skip-browser` declines it along with
agent-browser, and `off`/Camofox never use it.
- Installers regain --skip-computer-use / -SkipComputerUse (recorded as
`--without cua-driver`).
- The update message stops calling every default "browser tools".
- Refusal, non-zero exit and timeout now raise InstallError (pm/client.py
catches only InstallError; TimeoutExpired escaped entirely). The off-Windows
refusal names a real remedy instead of the generic 'retry'.
- stdin/stdout/stderr -> DEVNULL: with capture_output the stub's RunProgram
children inherit the pipes and hold run() open past the stub's exit (and
CPython re-communicates without a timeout after a kill). The stub prints
nothing under -y, so the output tail was dead code.
- TemporaryDirectory(ignore_cleanup_errors) as in Npm.unpack; unpack empties
its destination like extract(); drop redundant local imports.
- Tests: CompletedProcess over hand-rolled stubs, no raising=False, timeout
case added, change-detector pin-suffix test removed.
CI run 36189416163 failed after "git: unpacking": executing the cached
fetch-<sha> PortableGit PE in place left it handle-held (Defender
on-execute scan / the stub's RunProgram child chain) past pm's ~2 s
_remove_entry retry, so download cleanup raised WinError 32. Review
(teknium1, 5 threads) adds the rest:
- Git.unpack now copies the artifact into a .sfx-* dir beside the
staging tree and executes the copy; pm's .staging-* teardown
(ignore_errors) owns that path, so any hold lands on a disposable
path and the cache dir only ever holds read handles
- refuse off-Windows with the cross-host trade-off stated instead of a
raw PermissionError; documented in package-management.md and the PR
- the extractor is a GUI-subsystem stub that is silent under -y: error
messages now carry the exit code and the usual causes (disk full,
path length, antivirus) instead of promising captured output; the
docstring no longer claims "no GUI" and records that the stub shows
an Extracting window and runs the vendor post-install
- install.ps1: WaitForExit(600000) + Kill() mirrors pm's timeout=600,
Fail reports the exit code and the silence
- tests: the OS-refused-exec assertion and the not-in-text change
detectors are replaced by subprocess-argv invariants (scratch copy
location, exit-code message, off-Windows guard before any execution)
Related to #122512
(cherry picked from commit adaf76a286bebe30c89aee4174df27f1945b60c9)
Windows 10 boxes whose System32 tar.exe cannot run the bzip2 filter die
at stage=prerequisites with "unable to run program bzip2 -d" while
extracting the pinned Git-2.53.0.3 tar.bz2 (#122512). Repin git for
both win32 targets to git-for-windows' PortableGit self-extracting 7z,
which carries its own extractor and the bundled usr/bin/bash.exe:
- pm/lock.json: new artifact urls + sha256 (the pin authority)
- scripts/install.ps1: generated fragment regenerated; Get-PinnedGit
downloads the SFX and waits on it explicitly (the stub is a
GUI-subsystem exe, so PowerShell's & does not wait); the System32
tar invocation and its msys symlink excludes go away
- pm/packages.py: Git.fetch_url/Git.unpack run the self-extractor
after the sha256-verified download
- pm/store.py: drop the now-callerless git_msys branch of extract_tar
- tests: RED->GREEN test runs the real Get-PinnedGit against a
bzip2-less System32 tar.exe stub with no bzip2 on PATH; the three
obsolete tar-contract tests and the install.ps1/PM msys-links parity
test are replaced by a no-external-decompressor contract test
Closes#122512
(cherry picked from commit 4915304213495d3207ec6cd659e57cd16ef08d60)
BtbN deletes daily autobuild tags after 14 days and keeps the last build
of each month for two years (util/prunetags.sh). btbn_index took the
newest daily, so every `pm update ffmpeg` wrote a pin that 404s two weeks
later (#125350, #122240). It now skips the current month's tags (still
dailies the next build displaces) and every tag but the last in each
finished month, so the tool only ever pins a retained build.
The test replaced FFmpeg's gaps with only the bionic entry, which now
turns the musl targets back on and asks BtbN for a glibc build that
FFmpeg.fetch_url refuses on musl. Extend the real table instead.
Keep one PM invariant (a musl ELF userland resolves to linux-<arch>-musl
and the default closure is satisfiable with musl-native artifacts) and
one installer invariant (uv_bootstrap_target picks musl uv from a musl
ELF interpreter even when ldd says glibc). Drop the lock-content and
URL-shape change-detectors.
Skip glibc-linked FFmpeg on musl in both default install and update roots.
Select musl uv during the standalone shell bootstrap, and let the native
userland resolve libc before bootstrap Python build metadata.
Add regression checks for the closure, target precedence, and installer pins.
Trim the salvaged tests to one invariant: after a takeover rebuild, the
superseded generation built two days ago is gone. The update-completion
event-order asserts and the repair_dependencies `events == [sync,
collect]` test pin the call sequence, not the outcome; the update and
repair outcomes are covered live by
tests/e2e/core/upgrade/pm/test_generation_gc.py. The fake venv_sync in
the completion-process fixture keeps its collect_superseded_generations
stub so the module still imports.
Narrow the salvaged fix: keep store_root() returning $HERMES_HOME/tools as
spelled, and resolve the interpreter's directories only where
pm/runtime.py::_inputs hashes them. Resolving store_root() itself moves every
path built from store_root().parent (features.json, the native-build state dir,
writable_store_root's manifest probe), so after the update a symlinked tools
dir would read no feature selections and every install under a symlinked
parent would re-key.
Only directories resolve: the pinned CPython's bin/python3 is itself a symlink
to python3.X, so a full resolve() would re-key every install. A selected.json
written before this change (spelled path) is still accepted, so installs that
were current stay current; only per-task homes that were already re-staging
every launch re-stage once more and then converge.
Tests: the salvaged store_root alias tests are replaced by one invariant through
prepare_runtime (home, per-task home and real path share one generation; a
pre-fix record is not re-staged). Red on origin/main.
`dist` sat in the same every-depth exclusion set, so the generation
workspace lost the tracked `plugins/kanban/dashboard/dist/` and
`plugins/hermes-achievements/dashboard/dist/` bundles. The managed
environment's editable install runs from that snapshot, so
`get_bundled_plugins_dir()` resolves into it and
`dashboard_ui.py::serve_plugin_asset` 404s both tabs' entry bundles.
Keep `dist` excluded at the root (build output) and copy it below a
package root.
Refs #124075
Co-authored-by: Rafsanjani Castro Satria Chandra <330302617+oswaldvalois@users.noreply.github.com>
`_copy_core_inputs` listed `uv.lock` in the names its copytree `ignore`
drops, and copytree applies that callback at every depth, so the member
lock `pm/uv.lock` never reached the generation workspace.
`pm/runtime.py::_inputs()` hashes `<workspace>/pm/{pyproject.toml,uv.lock}`
to key the PM runtime, so every `hermes pm` command run by the managed
environment's `hermes` died at preflight:
FileNotFoundError: .../workspace/pm/uv.lock
Drop `uv.lock` from the exclusion set. The root lock is unaffected: the
root pass copies only the explicit `files` set (never `uv.lock`), and
`lock_and_sync` seeds or resolves the root lock itself.
Refs #124075
test_openssl_installs_once_and_rejects_damaged_shared_install and
test_native_build_command_preserves_failures_and_spaces each carried an
inline copy of the _powershell invocation with the pre-330ff28d44 30s
timeout. 330ff28d44 already established that a real powershell.exe child
on a cold CI runner can exceed 30s and raised the shared helper to 180s,
but these duplicated call sites kept the stale budget and timed out
intermittently on the windows-latest-32-arm-core lane.
Both now call _powershell(script, HELPER, tmp_path): identical argv and
environment construction (the inline setdefault calls were no-ops on
Windows), with the helper's documented cold-runner budget. No product
script changes.
The package-manager switch (#102765) made engine detection read only PM's
store and pinned every llama.cpp backend to b10362. A machine with an
engine under runtimes/llamacpp/b<tag>/<backend>/ reported no engine, so
the local models pane showed one-click setup with models already on disk.
installed_engine() now moves the newest pre-PM install into the store
once. The old manifest's archive digests become the PM identity, the
package verifier runs llama-server --version, and os.rename puts the
directory in place before facts.json records it. A store that already
holds an engine is left alone, a store lock held by another PM operation
defers the move, and an install that fails verification stays where it
is for the rest of the process.
The lock returns to b10964 on all five backends, the default before the
PM switch. With b10362 pinned, the pane offered a downgrade from the
moved engine as an update. Upstream renamed the ROCm archives to 10.0 at
b10767, and the hip assets follow. Only the CUDA build is measured at
b10964.
uv treats `foo_bar` and `foo-bar` as the same extra (PEP 685), so an
exact comparison would drop a recorded extra whose spelling differs from
its declaration. Compare normalized names; the recorded spelling is still
what reaches uv.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The venv ledger unions its recorded extras into every sync. When a source
update removes an extra (hindsight, 73c598e319/285768fdfb), installs that
recorded it pass `--extra hindsight` to `uv sync` forever and fail with
"Extra `hindsight` is not defined in any project's optional-dependencies
table". Every CLI launch then retries the failed source-update completion,
and gateways keep booting the previous generation's code.
Drop recorded extras the checkout's pyproject no longer declares, in both
the sync target and the currency probe. An explicitly requested unknown
extra still fails loudly.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sitting a plugin out after a fetch failure left the recorded stamp stale on
purpose, so every launch resynced, and the spawned-process guard in
prepare_launch raised "dependency sync left this install out of date".
One retry covers a blip; after that the plugin is disabled with the reason,
and hermes plugins enable restores it. Only requires_hermes still sits out,
which boot skips the same way.
enabled_member_dirs imported hermes_cli.plugins_manifest before its loop, so a
PM closure with no plugins selected needed the application's utils module
(tests/scripts/test_source_driver.py builds exactly that tree).
The requires_hermes sit-out test relied on the host's version identity. A
tagless CI checkout has no parseable version, which makes the gate permissive.
An update disabled any plugin that failed a trial build or its
requires_hermes check. Both can be about us, not the plugin: an untagged
source checkout reads as an older release (#122054), so requires_hermes
misjudges a fine plugin, and a download failure says nothing about the
plugin's code. Those now sit the plugin out of the build: config stays
untouched and it rejoins once the cause clears.
Disabling still happens on evidence about the plugin: requires-python vs
the pinned interpreter, manifest_version, an invalid declaration, a uv
resolution conflict, or its own build backend failing (new BuildFailure,
keyed on uv's 'The build backend returned an error').
enabled_member_dirs now skips a requires_hermes misfit instead of
raising. Boot's currency check raised on it before any sync could run, so
the launch path never reached the update sync. The loader skips such a
plugin anyway; admission still refuses enabling one.
Venv.apply refused the whole graph when any secondary profile's
config.yaml was unreadable, so one broken sibling config failed every
update. Update syncs now leave that profile's plugins out of the union,
report it (stderr + receipt warning), and build the rest; the profile's
plugins rejoin on the next sync once its config is fixed. Boot currency
already skips broken secondaries, so the result reads as current.
Ordinary syncs keep refusing to shrink the recorded graph.
An update resolves the enabled plugin union against the NEW core. A plugin
admitted against the old core can stop fitting when core moves (managed
Python 3.13 -> 3.14 vs a member's requires-python <3.14, a requires_hermes
upper bound, a bumped pin), and the whole update then died after the
source swap with a non-resolver InstallError whose 'retry' hint failed the
same way every time.
Update syncs now pass evict_incompatible_plugins=True (update completion,
historical takeover, launch-time completion, venv_sync, post-update
drift). PM screens statically first (requires-python vs the target
interpreter, manifest/requires_hermes), then, if the rest still fails,
builds core alone to prove the plugins are the cause and re-adds members
in config order, disabling each one that breaks the build. Misfits land in
plugins.disabled (memory.provider cleared) in every home that enables
them, published through the existing journaled change hook (the journal
now carries several configs), and are reported on stderr + receipt
warnings. Admission and ordinary syncs still refuse; only a core that
cannot build on its own fails an update.
A process that could not adopt a newly published dependency generation keeps
importing the old one. restart_needed() names that case (and stays silent for
dev venvs, Nix and anything not booted from a PM generation, so no false
restart prompts); adopt_selected() lets callers move onto the selection before
loading new code. The test helpers publish real generations and make the test
process run from one.
(cherry picked from commit 978abe8ec4b852778b8c63bcefdf141db51bc703)
(cherry picked from commit 28be27334adc531db59bf37a02a64acaaf47a2ee)
PR #122161 stopped boot from activating the in-tree venv or .venv when
PM has committed no environment. Six Linux tests and four Windows tests
still used that tree to supply their probe modules.
- Launcher tests put the probe in a committed generation. A custom
HERMES_HOME is its own dependency root, so it gets its own commit.
- The legacy row of test_pre_pm_base_dependencies_activate_only_at_boot
asserted the removed behavior. test_boot_never_activates_the_pre_pm_venv
now covers the inverse. The payload row stays.
- The mint payload fixture writes manifest.json as real payloads do, so
boot selects the payload venv.
- The PowerShell activate test expects PYTHONPATH to be the checkout
alone. The bootstrap .venv packages do not leak in.
Under the updater's claim, sync whenever dependencies are not current
rather than only when nothing is committed. The tail's own children
and post-sync verification children are current and stay no-ops. A
stale generation still committed from the previous Python pin is the
same ABI trap as the pre-PM venv, and it now syncs too. If the sync
still leaves the tree out of date, raise instead of relaunching into
another sync.
prepare_launch returned early for any process running under the
updater's own claim, so it would not re-run the completion tail. That
also covered processes the updater spawns before PM commits a
generation (a restarted gateway), which then booted with no
environment: previously on the pre-PM venv, now refused.
Under the updater's claim with nothing committed, sync the dependency
generation (carrying the legacy venv's extras, as the first sync
always has), skip the tail since that belongs to the updater, and
relaunch on the store Python. The relaunched process sees the commit
and returns early as before, so the no-recursion guard still holds.
With nothing committed, activate_dependencies fell back to the in-tree
venv/.venv. After an update that venv was built for the old interpreter
(uv CPython 3.11) while the process ran PM's store Python 3.14, so every
compiled module in it was unloadable: the messaging gateway's Group Chat
worker died on `No module named 'pydantic_core._pydantic_core'` until PM
committed a generation ~40 minutes later and deleted the old venv.
committed_venv() returns the committed generation or a sealed payload's
environment, never the in-tree venv. Boot activation and child
activation environments use it. With nothing committed, a venv/Nix
interpreter keeps its own packages; PM's bare store Python refuses with
the repair remedy instead of running on inherited paths. selected_venv
keeps its contract because pre-PM updaters import it after the swap.
uv identifies a workspace member by its declared project name, so the
same plugin enabled in two profiles declares one name twice and the
dependency sync fails with 'Two workspace members are both named ...'.
Metadata-only members (no build backend) now carry the unique member
key in their name, exactly like manifest-only members already do. A
buildable member keeps the name it declares, since uv verifies it
against the package metadata its backend produces.
The npm tarball ships every bin/agent-browser-* as 0644; agent-browser's
own postinstall sets the exec bit, and pm runs no postinstall. The first
browser_navigate auto-installs agent-browser and then fails with
PermissionError.
activate(allow_incomplete=True) discarded the venv verdict but still
computed it, and venv_is_current reads plugin selection through the
application config reader (ruamel). The update child runs the bare
bootstrap interpreter, so `hermes update` failed with "No module named
'ruamel'" once it published tools before the sync.
Browser tools find agent-browser only in PM's store or on PATH and their
readiness check never installs it, so a fresh install silently had no
browser_* tools. Package.default marks an optional package that a bare
`pm install` also carries; a failed download of it warns instead of
failing the install. `pm install --without NAME` records the opt-out in
declined-packages.json beside PM's install state, and naming the package
explicitly clears it. agent-browser and chromium gain a Termux gap: Termux
owns its browser stack and there is no bionic Chromium.