Commit Graph

258 Commits

Author SHA1 Message Date
teknium1
92ce6f4760 fix(pm): heal agent-browser entries staged without the exec bit
28edf7c75a made stage() chmod the native agent-browser binary, but
entries staged before it (lazy installs since the pm store landed) stay
0644 forever: every browser_* call fails with PermissionError and the
store never restages because the entry verifies. Restore the bit in
place when the entry is looked up on its own target (modes are not part
of the pinned digest); an entry that cannot be chmodded counts as not
installed so ensure() restages it.
2026-09-28 10:11:33 -07:00
JoaoMarcos44
0b105e0714 fix(pm): report unwritable install state once 2026-09-28 06:42:00 -07:00
Yuan Li
80e834b129 fix(pm): preserve transport knobs under no_config isolation, strip only index redirects
The isolation filter now drops only settings that redirect WHERE packages
resolve from (index URLs, find-links, strategy, index credentials); transport
knobs (UV_NATIVE_TLS / UV_INSECURE_HOST / UV_HTTP_TIMEOUT) survive — corporate
networks need them to reach the pinned wheel URLs at all.

Addresses the review note on the no_config scope in the #125071 triage.
2026-09-28 06:42:00 -07:00
Yuan Li
32aeb9ef7b fix(pm): honor no_config isolation for bridged index settings in uv runs
stage_runtime (the pinned pm/ self-repair builder) passes no_config=True,
but that only appended uv's --no-config, which removes config FILES. The
pip-mirror bridge (pm.index_config) reaches uv as env vars instead, so a
user's pip.conf index-url still rode into the runtime sync via
runtime_environment() -> _base_environment(env=None). Re-resolving the
official-index pm/uv.lock against a mirror changes the resolution, and
uv sync --locked refuses — breaking the top-level update / doctor /
pm-repair recovery paths on any machine with a pip mirror (#124418).

When no_config is set, strip forwarded index/transport knobs from the
child env: a fully-pinned graph never needs a mirror to resolve, and the
default path keeps the bridge so mirrored/air-gapped application installs
still work.
2026-09-28 06:42:00 -07:00
teknium1
dd3ba4b4b8 fix(pm): close the lease create→flock race on the writer side, no grace window
A time-based grace kept a just-abandoned lease 'held', so back-to-back `hermes pm repair` runs
never collected the generation the previous run execv'd out of (e2e test_generation_gc red).
Rename+replace is out (msvcrt byte locks refuse renames on Windows), so the writer now re-checks
its lease path after taking the flock and retakes when a peer's prune unlinked it in between; the
pruner only unlinks while holding the lock, so a visible path after the lock is always ours.
2026-09-28 04:44:29 -07:00
teknium1
6da0d5a887 fix(review): fail closed on unopenable or just-created runtime leases
Review majors on #126179 (runtime_state._prune_unlocked_leases):

1. The boot-time prune opened every sibling lease with O_RDWR and caught
   only FileNotFoundError, so one lease this user cannot open (a root-owned
   0o600 file from `sudo hermes` on the same checkout, the #125525 class)
   raised PermissionError inside activate_dependencies and hermes_bootstrap
   exited 1 with "run hermes pm repair" on every launch. Any other OSError
   now counts the lease as HELD: GC stays fail-closed, boot never aborts.

2. lease_directory creates (O_CREAT|O_EXCL) and flocks in two syscalls, and
   pm/worker.py, pm/launch.py and pm/environments.py call it outside
   runtime_lock, so a peer's prune in that window could take the
   non-blocking lock and unlink the file; the first process then held a
   lock on an unlinked inode and its pin was invisible for its lifetime
   (collect_generations could rmtree a running generation). A lease younger
   than LEASE_GRACE_SECONDS (60 s) is never a prune candidate and counts as
   held. Chosen over a temp-name + os.replace scheme because it is three
   lines, keeps a crash-between-create-and-flock leak collectable, and a
   60 s margin dwarfs a microsecond window.

test_next_reader_removes_lease_left_by_hard_exit ages the hard-exited
lease past the grace window before asserting the next reader removes it;
its invariant is unchanged.
2026-09-28 04:44:29 -07:00
chelsealong
42941c0a37 fix(pm): put the checkout launcher ahead of the venv's own hermes script (#124627, salvage #124828)
The runtime venv installs the core project editable against the PM
build snapshot (workspace/), so venv/bin/hermes always imports that
frozen copy. PM only rebuilds the venv on lock/extras/python-pin/plugin
changes, so a source-only update leaves the copy stale while the
gateway keeps running the checkout via activate_dependencies's
sys.path override. Any child that resolves `hermes` off PATH (terminal
tool, cron, kanban workers) still finds venv/bin first and runs the
stale snapshot instead — including `gateway install`/`restart`, which
then pins launchd/systemd to the snapshot's launcher path.

Prepend the checkout's own .hermes/bin ahead of venv/bin in
activate_dependencies so `hermes`/`hermes-acp` always resolve to the
launcher that puts the checkout first, regardless of what else lives
in the venv. This also fixes tools/environments/local.py's
_resolve_hermes_bin_dir(), whose shutil.which("hermes") now finds the
launcher dir instead of the venv's console script, so no separate guard
on PROJECT_ROOT in generate_launchd_plist / systemd unit generation is
needed: the snapshot's hermes_cli is never the running one.

Fixes #124627
2026-09-28 04:44:29 -07:00
Ahmett101
b23d4f4b32 fix(pm): collect abandoned runtime leases (#125609, salvage #125632)
lease_directory releases its per-invocation lease file only through atexit,
which os.execv (hermes_bootstrap's re-exec into the managed environment) and
os._exit (gateway shutdown_watchdog._hard_exit, hermes_cli/main.py past
finalization) never run: a 4s supervisor restart loop left ~860 lease files an
hour in the SELECTED generation, which generation GC never visits.

The kernel lock dies with the process even when atexit does not, so the file
itself is the only leak: every reader (the next lease_directory and
leases_held) now unlinks lease files nobody holds a lock on. Both run under
runtime_lock at boot / in the collector, so a lease between O_CREAT|O_EXCL and
its lock is never pruned from under a live process.
2026-09-28 04:44:29 -07:00
teknium1
a31e618e70 fix(pm): download npm-hosted tools through the user's npm registry
pm/lock.json and the Npm/AgentBrowser templates record registry.npmjs.org, and
pinned_source handed that URL straight to the downloader, so ~/.npmrc /
npm_config_registry were ignored and closed networks could not provision npm or
agent-browser. pinned_source now routes public-npm URLs through the configured
registry (npm's own precedence: npm_config_registry, then the user npmrc); the
lock's SHA256 still verifies the bytes and progress keeps the lockfile URL.
npm_dist_tags reads from the same registry.

The E2E cell test_npm_registry_mirror_serves_pm_npm_download drops its gate.

Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
2026-09-28 02:00:41 -07:00
teknium1
e0fae5566a fix: Hermes's own lint, activate, tui and MCP launchers use PM's node/npx/uv, never the user's
Post-write .js/.ts lint on the local backend runs PM's node/npx instead of
the terminal shell's PATH (user dirs first). pm.activate always puts the
store dirs in front, also on drift, activating only packages whose whole
chain is installed. tui_gateway puts the store dirs ahead of ~/.local/bin.
Bare npx/npm/node/uv/uvx MCP commands resolve to PM's copies; the system
fallback tables are gone and an absolute command: stays the user's.
2026-09-27 22:30:55 -07:00
teknium1
c274b9e9c9 test: Windows PATH .cmd preference applies to non-PM tools only 2026-09-27 22:04:26 -07:00
teknium1
b63c138d78 fix: Hermes never falls back to the user's node/npm/npx/uv
Maintainer ruling: only Hermes and only its packaged package managers
(uv/pip/node/npm) are ever used; no PATH fallback when the managed tool is
missing, no "prefer the user's if new enough".

- hermes_constants.find_node_executable: node/npm/npx resolve to PM's
  installed copy or None. Every caller already pm.ensure()s on None, so a
  missing runtime is now provisioned instead of silently borrowing the
  user's Node (native-addon ABI / npm cache mismatches).
- agent/lsp/install._install_npm: pm.ensure('npm') when PM npm is absent,
  instead of failing over to whatever npm is on PATH.
- gateway._append_node_dir_for_service: stop baking the invoker's PATH node
  dir into generated systemd/launchd units.
- main_install_repair._resolve_node_runtime_npm: drop the PATH re-scan for
  another npm.
- source_build.source_product_current: run the freshness reader only with
  PM's node.
- doctor: Node/npm rows and npm audit use PM's copies (Termux APT distro
  keeps its system Node).
- install.sh ensure_uv / install.ps1 Get-Uv: always stage the pinned uv
  artifact; delete the "uv on PATH if new enough" developer shortcut.
2026-09-27 22:04:26 -07:00
teknium1
27062c3474 fix: install cua-driver and the Browser Use CLI by default again
Computer use and browser use are meant to work out of the box. The PM rewrite
(3d12e86ef1) and the MSIX installer rework (47f4ab3a17) dropped the
install-time cua-driver fetch (7060ac7bed) and the Browser Use CLI install
(baa6b2e34d). On a fresh install the computer_use check_fn therefore stayed
False, so the tool never reached the model and its lazy ensure could not fire,
and browser_exec quietly fell back to the built-in tools.

- cua-driver is a default PM package, so the installers, a bare
  `hermes pm install` and `hermes update` carry it on every target it builds
  for. Adds the missing Android gap (the lock has no bionic artifact).
- The shared default-tool step, which the installers (via source completion)
  and `hermes update` both run, provisions the Browser Use CLI for the default
  and explicit Browser Use backends. `--skip-browser` declines it along with
  agent-browser, and `off`/Camofox never use it.
- Installers regain --skip-computer-use / -SkipComputerUse (recorded as
  `--without cua-driver`).
- The update message stops calling every default "browser tools".
2026-09-27 19:18:58 -07:00
kshitijk4poor
084645fba4 fix(pm): PortableGit unpack raises InstallError and never pipes the extractor
- Refusal, non-zero exit and timeout now raise InstallError (pm/client.py
  catches only InstallError; TimeoutExpired escaped entirely). The off-Windows
  refusal names a real remedy instead of the generic 'retry'.
- stdin/stdout/stderr -> DEVNULL: with capture_output the stub's RunProgram
  children inherit the pipes and hold run() open past the stub's exit (and
  CPython re-communicates without a timeout after a kill). The stub prints
  nothing under -y, so the output tail was dead code.
- TemporaryDirectory(ignore_cleanup_errors) as in Npm.unpack; unpack empties
  its destination like extract(); drop redundant local imports.
- Tests: CompletedProcess over hand-rolled stubs, no raising=False, timeout
  case added, change-detector pin-suffix test removed.
2026-09-28 02:04:08 +05:30
finn763
7a40c28750 fix(pm): execute the pinned git SFX from a scratch copy, not the cache
CI run 36189416163 failed after "git: unpacking": executing the cached
fetch-<sha> PortableGit PE in place left it handle-held (Defender
on-execute scan / the stub's RunProgram child chain) past pm's ~2 s
_remove_entry retry, so download cleanup raised WinError 32. Review
(teknium1, 5 threads) adds the rest:

- Git.unpack now copies the artifact into a .sfx-* dir beside the
  staging tree and executes the copy; pm's .staging-* teardown
  (ignore_errors) owns that path, so any hold lands on a disposable
  path and the cache dir only ever holds read handles
- refuse off-Windows with the cross-host trade-off stated instead of a
  raw PermissionError; documented in package-management.md and the PR
- the extractor is a GUI-subsystem stub that is silent under -y: error
  messages now carry the exit code and the usual causes (disk full,
  path length, antivirus) instead of promising captured output; the
  docstring no longer claims "no GUI" and records that the stub shows
  an Extracting window and runs the vendor post-install
- install.ps1: WaitForExit(600000) + Kill() mirrors pm's timeout=600,
  Fail reports the exit code and the silence
- tests: the OS-refused-exec assertion and the not-in-text change
  detectors are replaced by subprocess-argv invariants (scratch copy
  location, exit-code message, off-Windows guard before any execution)

Related to #122512

(cherry picked from commit adaf76a286bebe30c89aee4174df27f1945b60c9)
2026-09-28 02:04:08 +05:30
finn763
6d9ae2da6f fix(install): stage pinned git without tar/bzip2 (PortableGit SFX)
Windows 10 boxes whose System32 tar.exe cannot run the bzip2 filter die
at stage=prerequisites with "unable to run program bzip2 -d" while
extracting the pinned Git-2.53.0.3 tar.bz2 (#122512). Repin git for
both win32 targets to git-for-windows' PortableGit self-extracting 7z,
which carries its own extractor and the bundled usr/bin/bash.exe:

- pm/lock.json: new artifact urls + sha256 (the pin authority)
- scripts/install.ps1: generated fragment regenerated; Get-PinnedGit
  downloads the SFX and waits on it explicitly (the stub is a
  GUI-subsystem exe, so PowerShell's & does not wait); the System32
  tar invocation and its msys symlink excludes go away
- pm/packages.py: Git.fetch_url/Git.unpack run the self-extractor
  after the sha256-verified download
- pm/store.py: drop the now-callerless git_msys branch of extract_tar
- tests: RED->GREEN test runs the real Get-PinnedGit against a
  bzip2-less System32 tar.exe stub with no bzip2 on PATH; the three
  obsolete tar-contract tests and the install.ps1/PM msys-links parity
  test are replaced by a no-external-decompressor contract test

Closes #122512

(cherry picked from commit 4915304213495d3207ec6cd659e57cd16ef08d60)
2026-09-28 02:04:08 +05:30
kshitijk4poor
7546680ff8 fix(pm): pm update pins only BtbN month-end ffmpeg builds
BtbN deletes daily autobuild tags after 14 days and keeps the last build
of each month for two years (util/prunetags.sh). btbn_index took the
newest daily, so every `pm update ffmpeg` wrote a pin that 404s two weeks
later (#125350, #122240). It now skips the current month's tags (still
dailies the next build displaces) and every tag but the last in each
finished month, so the tool only ever pins a retained build.
2026-09-28 00:51:32 +05:30
teknium1
5cbca02510 test(pm): keep FFmpeg's musl gaps when the re-pin test narrows its gap table
The test replaced FFmpeg's gaps with only the bionic entry, which now
turns the musl targets back on and asks BtbN for a glibc build that
FFmpeg.fetch_url refuses on musl. Extend the real table instead.
2026-09-27 03:24:39 -07:00
teknium1
f5c3c7d70a test: trim musl coverage to two invariants
Keep one PM invariant (a musl ELF userland resolves to linux-<arch>-musl
and the default closure is satisfiable with musl-native artifacts) and
one installer invariant (uv_bootstrap_target picks musl uv from a musl
ELF interpreter even when ldd says glibc). Drop the lock-content and
URL-shape change-detectors.
2026-09-27 03:24:39 -07:00
JoaoMarcos44
24487f3db6 fix(pm): keep musl installs on compatible runtime artifacts
Skip glibc-linked FFmpeg on musl in both default install and update roots.
Select musl uv during the standalone shell bootstrap, and let the native
userland resolve libc before bootstrap Python build metadata.

Add regression checks for the closure, target precedence, and installer pins.
2026-09-27 03:24:39 -07:00
JoaoMarcos44
8371dd17f6 fix(pm): select native musl artifacts on Linux 2026-09-27 03:24:39 -07:00
teknium1
ccff8eb3ce test(pm): keep the takeover generation assert; drop call-order checks
Trim the salvaged tests to one invariant: after a takeover rebuild, the
superseded generation built two days ago is gone. The update-completion
event-order asserts and the repair_dependencies `events == [sync,
collect]` test pin the call sequence, not the outcome; the update and
repair outcomes are covered live by
tests/e2e/core/upgrade/pm/test_generation_gc.py. The fake venv_sync in
the completion-process fixture keeps its collect_superseded_generations
stub so the module still imports.
2026-09-27 02:55:39 -07:00
JoaoMarcos44
c11a6bc0b2 fix(pm): collect generations after maintenance syncs 2026-09-27 02:55:39 -07:00
teknium1
6a17705227 fix(pm): canonicalize the store only in the PM runtime identity
Narrow the salvaged fix: keep store_root() returning $HERMES_HOME/tools as
spelled, and resolve the interpreter's directories only where
pm/runtime.py::_inputs hashes them. Resolving store_root() itself moves every
path built from store_root().parent (features.json, the native-build state dir,
writable_store_root's manifest probe), so after the update a symlinked tools
dir would read no feature selections and every install under a symlinked
parent would re-key.

Only directories resolve: the pinned CPython's bin/python3 is itself a symlink
to python3.X, so a full resolve() would re-key every install. A selected.json
written before this change (spelled path) is still accepted, so installs that
were current stay current; only per-task homes that were already re-staging
every launch re-stage once more and then converge.

Tests: the salvaged store_root alias tests are replaced by one invariant through
prepare_runtime (home, per-task home and real path share one generation; a
pre-fix record is not re-staged). Red on origin/main.
2026-09-27 02:55:05 -07:00
funky-xamarin
f7317b344f test(pm): prove runtime alias reuse retains a real lease 2026-09-27 02:55:05 -07:00
funky-xamarin
42b1f19887 fix(pm): canonicalize default store before runtime identity 2026-09-27 02:55:05 -07:00
teknium1
9fa34cfb28 fix(pm): keep bundled plugins' dashboard dist/ in the build snapshot
`dist` sat in the same every-depth exclusion set, so the generation
workspace lost the tracked `plugins/kanban/dashboard/dist/` and
`plugins/hermes-achievements/dashboard/dist/` bundles. The managed
environment's editable install runs from that snapshot, so
`get_bundled_plugins_dir()` resolves into it and
`dashboard_ui.py::serve_plugin_asset` 404s both tabs' entry bundles.

Keep `dist` excluded at the root (build output) and copy it below a
package root.

Refs #124075

Co-authored-by: Rafsanjani Castro Satria Chandra <330302617+oswaldvalois@users.noreply.github.com>
2026-09-27 02:54:35 -07:00
Daniel Bader
7523696fb6 fix(pm): keep a workspace member's own uv.lock in the build snapshot
`_copy_core_inputs` listed `uv.lock` in the names its copytree `ignore`
drops, and copytree applies that callback at every depth, so the member
lock `pm/uv.lock` never reached the generation workspace.
`pm/runtime.py::_inputs()` hashes `<workspace>/pm/{pyproject.toml,uv.lock}`
to key the PM runtime, so every `hermes pm` command run by the managed
environment's `hermes` died at preflight:

    FileNotFoundError: .../workspace/pm/uv.lock

Drop `uv.lock` from the exclusion set. The root lock is unaffected: the
root pass copies only the explicit `files` set (never `uv.lock`), and
`lock_and_sync` seeds or resolves the root lock itself.

Refs #124075
2026-09-27 02:54:35 -07:00
Hermes Agent
5838f13725 test(pm): route build-deps PowerShell checks through the cold-runner helper
test_openssl_installs_once_and_rejects_damaged_shared_install and
test_native_build_command_preserves_failures_and_spaces each carried an
inline copy of the _powershell invocation with the pre-330ff28d44 30s
timeout. 330ff28d44 already established that a real powershell.exe child
on a cold CI runner can exceed 30s and raised the shared helper to 180s,
but these duplicated call sites kept the stale budget and timed out
intermittently on the windows-latest-32-arm-core lane.

Both now call _powershell(script, HELPER, tmp_path): identical argv and
environment construction (the inline setdefault calls were no-ops on
Windows), with the helper's documented cold-runner budget. No product
script changes.
2026-09-26 16:50:13 -05:00
emozilla
9c1ef17550 fix(local-runtime): move a pre-PM llama.cpp engine into the PM store and pin b10964
The package-manager switch (#102765) made engine detection read only PM's
store and pinned every llama.cpp backend to b10362. A machine with an
engine under runtimes/llamacpp/b<tag>/<backend>/ reported no engine, so
the local models pane showed one-click setup with models already on disk.

installed_engine() now moves the newest pre-PM install into the store
once. The old manifest's archive digests become the PM identity, the
package verifier runs llama-server --version, and os.rename puts the
directory in place before facts.json records it. A store that already
holds an engine is left alone, a store lock held by another PM operation
defers the move, and an install that fails verification stays where it
is for the rest of the process.

The lock returns to b10964 on all five backends, the default before the
PM switch. With b10362 pinned, the pane offered a downgrade from the
moved engine as an update. Upstream renamed the ROCm archives to 10.0 at
b10767, and the hip assets follow. Only the CUDA build is measured at
b10964.
2026-09-25 22:07:48 -04:00
Matt Healey
ac3ebbfedc fix(pm): match recorded extras to declarations by normalized name
uv treats `foo_bar` and `foo-bar` as the same extra (PEP 685), so an
exact comparison would drop a recorded extra whose spelling differs from
its declaration. Compare normalized names; the recorded spelling is still
what reaches uv.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 13:48:40 -04:00
Matt Healey
24e8c341c1 fix(pm): prune recorded extras the tree no longer declares
The venv ledger unions its recorded extras into every sync. When a source
update removes an extra (hindsight, 73c598e319/285768fdfb), installs that
recorded it pass `--extra hindsight` to `uv sync` forever and fail with
"Extra `hindsight` is not defined in any project's optional-dependencies
table". Every CLI launch then retries the failed source-update completion,
and gateways keep booting the previous generation's code.

Drop recorded extras the checkout's pyproject no longer declares, in both
the sync target and the currency probe. An explicitly requested unknown
extra still fails loudly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 13:48:40 -04:00
ethernet
067934a106 Merge pull request #122093 from benbarclay/fix/pm-agent-browser-exec-bit
fix(pm): make the staged agent-browser binary executable
2026-09-25 00:35:56 -04:00
ethernet
2ef41d2b58 Merge pull request #122103 from NousResearch/ethie/pm-evict-incompatible-plugins
fix(pm): hermes update disables plugins that no longer fit instead of failing
2026-09-25 00:08:10 -04:00
ethernet
e71f4dd4ac Merge pull request #122221 from NousResearch/fix/anthropic-lazy-install
Opt-in provider SDKs (anthropic, bedrock) install and swap into the running process on first use
2026-09-25 00:05:51 -04:00
ethernet
0978aca962 fix(pm): retry a plugin fetch failure once, then disable it
Sitting a plugin out after a fetch failure left the recorded stamp stale on
purpose, so every launch resynced, and the spawned-process guard in
prepare_launch raised "dependency sync left this install out of date".
One retry covers a blip; after that the plugin is disabled with the reason,
and hermes plugins enable restores it. Only requires_hermes still sits out,
which boot skips the same way.
2026-09-25 00:00:09 -04:00
ethernet
10d066ffd8 fix(pm): keep the manifest import per plugin; pin the version in the sit-out test
enabled_member_dirs imported hermes_cli.plugins_manifest before its loop, so a
PM closure with no plugins selected needed the application's utils module
(tests/scripts/test_source_driver.py builds exactly that tree).

The requires_hermes sit-out test relied on the host's version identity. A
tagless CI checkout has no parseable version, which makes the gate permissive.
2026-09-25 00:00:09 -04:00
ethernet
c500ee8771 fix(pm): disable a plugin only on evidence about the plugin
An update disabled any plugin that failed a trial build or its
requires_hermes check. Both can be about us, not the plugin: an untagged
source checkout reads as an older release (#122054), so requires_hermes
misjudges a fine plugin, and a download failure says nothing about the
plugin's code. Those now sit the plugin out of the build: config stays
untouched and it rejoins once the cause clears.

Disabling still happens on evidence about the plugin: requires-python vs
the pinned interpreter, manifest_version, an invalid declaration, a uv
resolution conflict, or its own build backend failing (new BuildFailure,
keyed on uv's 'The build backend returned an error').

enabled_member_dirs now skips a requires_hermes misfit instead of
raising. Boot's currency check raised on it before any sync could run, so
the launch path never reached the update sync. The loader skips such a
plugin anyway; admission still refuses enabling one.
2026-09-25 00:00:09 -04:00
ethernet
cae75f0ead fix(pm): an unreadable secondary profile config cannot fail an update
Venv.apply refused the whole graph when any secondary profile's
config.yaml was unreadable, so one broken sibling config failed every
update. Update syncs now leave that profile's plugins out of the union,
report it (stderr + receipt warning), and build the rest; the profile's
plugins rejoin on the next sync once its config is fixed. Boot currency
already skips broken secondaries, so the result reads as current.
Ordinary syncs keep refusing to shrink the recorded graph.
2026-09-25 00:00:09 -04:00
ethernet
ccf631701a fix(pm): disable plugins that no longer fit instead of failing the update
An update resolves the enabled plugin union against the NEW core. A plugin
admitted against the old core can stop fitting when core moves (managed
Python 3.13 -> 3.14 vs a member's requires-python <3.14, a requires_hermes
upper bound, a bumped pin), and the whole update then died after the
source swap with a non-resolver InstallError whose 'retry' hint failed the
same way every time.

Update syncs now pass evict_incompatible_plugins=True (update completion,
historical takeover, launch-time completion, venv_sync, post-update
drift). PM screens statically first (requires-python vs the target
interpreter, manifest/requires_hermes), then, if the rest still fails,
builds core alone to prove the plugins are the cause and re-adds members
in config order, disabling each one that breaks the build. Misfits land in
plugins.disabled (memory.provider cleared) in every home that enables
them, published through the existing journaled change hook (the journal
now carries several configs), and are reported on stderr + receipt
warnings. Admission and ordinary syncs still refuse; only a core that
cannot build on its own fails an update.
2026-09-25 00:00:09 -04:00
ethernet
5d39ddd28b feat(pm): say when this process must restart to load the selected generation
A process that could not adopt a newly published dependency generation keeps
importing the old one. restart_needed() names that case (and stays silent for
dev venvs, Nix and anything not booted from a PM generation, so no false
restart prompts); adopt_selected() lets callers move onto the selection before
loading new code. The test helpers publish real generations and make the test
process run from one.

(cherry picked from commit 978abe8ec4b852778b8c63bcefdf141db51bc703)
(cherry picked from commit 28be27334adc531db59bf37a02a64acaaf47a2ee)
2026-09-24 23:40:17 -04:00
ethernet
e3d43c3c24 test: stop relying on the in-tree venv after #122161
PR #122161 stopped boot from activating the in-tree venv or .venv when
PM has committed no environment. Six Linux tests and four Windows tests
still used that tree to supply their probe modules.

- Launcher tests put the probe in a committed generation. A custom
  HERMES_HOME is its own dependency root, so it gets its own commit.
- The legacy row of test_pre_pm_base_dependencies_activate_only_at_boot
  asserted the removed behavior. test_boot_never_activates_the_pre_pm_venv
  now covers the inverse. The payload row stays.
- The mint payload fixture writes manifest.json as real payloads do, so
  boot selects the payload venv.
- The PowerShell activate test expects PYTHONPATH to be the checkout
  alone. The bootstrap .venv packages do not leak in.
2026-09-24 23:40:03 -04:00
ethernet
6ce9ed4223 fix(update): key the early-spawn sync on currency, not on a commit
Under the updater's claim, sync whenever dependencies are not current
rather than only when nothing is committed. The tail's own children
and post-sync verification children are current and stay no-ops. A
stale generation still committed from the previous Python pin is the
same ABI trap as the pre-PM venv, and it now syncs too. If the sync
still leaves the tree out of date, raise instead of relaunching into
another sync.
2026-09-24 22:56:59 -04:00
ethernet
3a42f0fe10 fix(update): commit dependencies for processes the update spawns early
prepare_launch returned early for any process running under the
updater's own claim, so it would not re-run the completion tail. That
also covered processes the updater spawns before PM commits a
generation (a restarted gateway), which then booted with no
environment: previously on the pre-PM venv, now refused.

Under the updater's claim with nothing committed, sync the dependency
generation (carrying the legacy venv's extras, as the first sync
always has), skip the tail since that belongs to the updater, and
relaunch on the store Python. The relaunched process sees the commit
and returns early as before, so the no-recursion guard still holds.
2026-09-24 22:53:06 -04:00
ethernet
c821ecdd8f fix(pm): never activate the pre-PM in-tree venv
With nothing committed, activate_dependencies fell back to the in-tree
venv/.venv. After an update that venv was built for the old interpreter
(uv CPython 3.11) while the process ran PM's store Python 3.14, so every
compiled module in it was unloadable: the messaging gateway's Group Chat
worker died on `No module named 'pydantic_core._pydantic_core'` until PM
committed a generation ~40 minutes later and deleted the old venv.

committed_venv() returns the committed generation or a sealed payload's
environment, never the in-tree venv. Boot activation and child
activation environments use it. With nothing committed, a venv/Nix
interpreter keeps its own packages; PM's bare store Python refuses with
the repair remedy instead of running on inherited paths. selected_venv
keeps its contract because pre-PM updaters import it after the swap.
2026-09-24 22:48:35 -04:00
fangliquan
425c5c1167 test(pm): exclude catalog Hindsight from legacy extra selection 2026-09-24 22:12:39 -04:00
liuhao1024
0fc90369a5 fix(pm): name virtual workspace members by their unique key
uv identifies a workspace member by its declared project name, so the
same plugin enabled in two profiles declares one name twice and the
dependency sync fails with 'Two workspace members are both named ...'.
Metadata-only members (no build backend) now carry the unique member
key in their name, exactly like manifest-only members already do. A
buildable member keeps the name it declares, since uv verifies it
against the package metadata its backend produces.
2026-09-25 09:24:18 +08:00
Ben Barclay
28edf7c75a fix(pm): make the staged agent-browser binary executable
The npm tarball ships every bin/agent-browser-* as 0644; agent-browser's
own postinstall sets the exec bit, and pm runs no postinstall. The first
browser_navigate auto-installs agent-browser and then fails with
PermissionError.
2026-09-25 11:19:18 +10:00
ethernet
7581ef4866 fix(pm): keep the pre-sync tool gate out of the venv check
activate(allow_incomplete=True) discarded the venv verdict but still
computed it, and venv_is_current reads plugin selection through the
application config reader (ruamel). The update child runs the bare
bootstrap interpreter, so `hermes update` failed with "No module named
'ruamel'" once it published tools before the sync.
2026-09-24 17:56:48 -04:00
ethernet
3addc73a83 pm: install agent-browser by default, with a recorded --without opt-out
Browser tools find agent-browser only in PM's store or on PATH and their
readiness check never installs it, so a fresh install silently had no
browser_* tools. Package.default marks an optional package that a bare
`pm install` also carries; a failed download of it warns instead of
failing the install. `pm install --without NAME` records the opt-out in
declined-packages.json beside PM's install state, and naming the package
explicitly clears it. agent-browser and chromium gain a Termux gap: Termux
owns its browser stack and there is no bionic Chromium.
2026-09-24 17:27:45 -04:00