Commit Graph

45142 Commits

Author SHA1 Message Date
Brooklyn Nicholson
69ccfa69e4 fix(desktop): map behind: -1 to null instead of clamping the sentinel to 0
The backend update-check endpoint answers behind: -1
(source_check.UPDATE_AVAILABLE_NO_COUNT) when the checkout is behind but
the count can't be computed — a shallow clone without a merge-base or an
unusable GitHub compare. mapBackendCheck clamped it to 0, making that
state byte-identical to the genuinely-up-to-date answer while
update_available still pitched the install, so every behind-based branch
disagreed with the overlay on the same screen and the changelog rendered
"what changed" over zero rows.

DesktopUpdateStatus already types this state as null ("never render it
as a literal number"), so pass the sentinel through as null. The
check-failed predicate is unchanged: it still keys on behind === null
with can_apply, which -1 correctly does not trigger (the check did run).

Also updates the pre-existing updates.test.ts expectation that encoded
the old clamp, and pins the sentinel in updates-backend-check.test.ts.
2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
31a68e9235 fix(tools): resolve the live plugin catalog once per plugins.list
The #119975 display-name lookup called get_live_catalog_entry() inside
the per-plugin loop of _plugin_rows, paying a full catalog resolution
(load_catalog_live: fetch or cache read + parse of every in-tree catalog
yaml, no memoization) once per installed plugin. plugins.manage list went
from O(1) to O(installed plugins) catalog resolutions, and a dead
catalog host cost one request timeout per plugin — the exact per-candidate
cost resolved_removed_entries() exists to eliminate.

Hoist the resolution next to pins/versions: catalog_titles() builds
{catalog_name: title} in one resolution and _plugin_server_rows reads
from the pre-resolved map, mirroring catalog_pins/catalog_versions.
Regression test counts load_catalog_live calls across a 3-plugin
listing: 3 before, 1 after.
2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
76e2111c14 fix(update): drop the cached live plugin catalog on boot after a code change
Bundled/sealed app updates never run hermes update's maintenance tail, so
their pre-update catalog snapshot stayed authoritative for the cache TTL.
Add a boot home step (once per installed revision, per profile) that drops
the active home's cache (#119340).
2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
f4a86aa135 fix(lint): sort named imports in screen-hero 2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
0f15d80a02 fix(contracts): regenerate gateway contract for hermes_not_connected enum 2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
f5fad3cedb chore: map contributor emails 2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
252402dfa2 fix(desktop): Screen on a managed Cloud backend says the release lacks it, not 'update the bot's Hermes'
Fresh fix for #120852 (PR #120860 was deleted; nothing to salvage).

Any display.status -32601 rendered 'Screen needs a newer Hermes / Update
the bot's Hermes to use Screen' on every surface, with no distinction
between a stale git install the user can update and a Portal-managed
release they cannot: against Hermes Cloud the managed tab reports 'latest
release / Up to date' at the same time, which reads as a contradiction.

The roster already knows which kind of backend a bot runs on: the registry
stamps connectionKind: 'cloud' onto the row (annotateBotSource; Connection
Kind 'cloud' is remote-shaped but kept distinct for exactly this kind of
difference). Add isManagedBackend(bot) and pick the copy accordingly:
managed backends get 'Screen is not available on this managed Hermes
release yet' (portalUnavailableManaged, localized in en/ja/zh/zh-hant);
self-upgradable backends keep the update instruction.

Applies to all three surfaces that render the unavailable state: the
Screen pane's empty state, the sidebar portal subtitle, and the hero
caption. The managed release genuinely lacks display.* (methods_display.py
is absent at v2026.9.21), so hiding the surface stays correct — only the
sentence was wrong.

Fixes #120852
2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
48cf36fb69 fix(tools): say 'MCP connection missing' when only the MCP connection is missing
Fresh fix for #119975 (PR #119993 was deleted; nothing to salvage).

Three paths collapsed into app_not_running with the sentence '<slug> is
not running. Start <slug>': (1) a server_json probe with the app running
AND its endpoint PRESENT — the exact case from the report, where the only
thing missing is Hermes' own MCP connection; (2) every non-server_json,
non-interactive_session liveness kind; (3) static and unregistered
liveness, which cannot observe the app at all yet still claimed it was
stopped. Telling the user to start an app that IS running is the wrong
instruction.

Add a hermes_not_connected LivenessState ('<app>'s MCP connection is
missing. Reconnect <app> in Hermes, then try again.'), map the
running+endpoint-present branch and the static/unknown fallback to it, and
wire it through the TUI gateway contract (PluginServerState) and the
Desktop Plugins tab (AgentPluginServerState, SERVER_TONE, serverStates
i18n in en/de/es/fr).

Also stop composing the sentence from the declaration's slug: describe()
takes an optional display_name, and _plugin_server_rows passes the
curated catalog title (fallback: the manifest name) so the Plugins tab
reads 'NVIDIA App' instead of a raw server slug.

Fixes #119975
2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
97e0bae47c fix(update): drop the cached live plugin catalog after an update
Salvages the intent of #119354 by JoaoMarcos44 (its _invalidate_update_cache
hook no longer exists on main — the update pipeline was rewritten around
source_completion/update_finish; the invalidation is re-landed at the new
post-update maintenance seam).

A plugin added to plugin-catalog/ stayed "not in the plugin catalog" for up
to 6h after hermes update: LIVE_CATALOG_TTL_SECONDS keeps the on-disk live
snapshot authoritative for 6h, the update pipeline never touched it, and
load_catalog_live() iterates only the live entries — so the stale pre-update
snapshot out-voted the newer in-tree catalog the update just installed
(#119340).

Drop HERMES_HOME/cache/plugin-catalog.json under the active home AND every
sibling profile's (the checkout is shared, mirroring the post-update state.db
guard's home + siblings sweep) in _run_post_update_maintenance. The next
fetch_live_catalog() re-fetches the published doc, or falls back to the
in-tree catalog while the network is down — both newer than what was deleted.

Fixes #119340 (fix A)
2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
680ea95123 fix(desktop): count apps/shared commits in the bundle-skew detector
Salvaged from #120253 by jonpol01 (Sweep) — the diff applies unchanged
onto current main.

RUNTIME_PATHS listed only apps/desktop/*, but apps/shared/src is compiled
into BOTH bundles: the renderer through the vite alias '@hermes/shared' ->
../shared/src (plus the file: dependency on apps/shared/package.json), and
the Electron main process by relative import (bootstrap-runner.ts,
hardening.ts, translucency.ts import '../../shared/src/...'). A fix
confined to shared code (the JSON-RPC layer, the gateway client) left the
installed app just as stale as a renderer change, but the skew detector
stayed silent about it.

Add apps/shared/src and apps/shared/package.json to RUNTIME_PATHS so
shared-only commits count toward desktopCommitsBehind/outOfSync; non-runtime
shared content (e.g. README.md) stays excluded, covered by the new
it.each regression test.

Fixes #120252
2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
fa65609351 fix(desktop): report a failed backend update check instead of "up to date"
Salvaged from #119803 by Overview3833 (rebased onto current main; the
diff itself applies unchanged).

`mapBackendCheck` folded the backend's `behind: null` into `behind: 0` and
returned a status without `error`. But `null` is the endpoint's "the check
could not run" answer — GitHub unreachable, rate limited, offline — and it
carries the explanation in `message`.

The remote overlay therefore fell through to the all-set copy ("the backend is
on the latest version") whenever the check failed, and the message the backend
sent was never shown. The local check path already reports `error:
'check-failed'` in this situation; only the backend mapping dropped it.

Map a null from a supported (git) backend to the same failure state, so the
overlay's `status.error` branch shows the backend's message with its retry
affordance (and the "update available" toast stays suppressed, since we don't
know the real distance). A check that ran and found no gap still reports
"latest", and backends that cannot self-update keep taking the unsupported
branch.

Fixes #119801
2026-09-27 06:26:52 -05:00
Brooklyn Nicholson
89cc403b23 fix(desktop): resolve the pre-launch home through the shared resolver
configuredElectronFlags() inlined HERMES_HOME normalization, so it missed
the two cases resolveDesktopHermesHome() handles: HERMES_DATA_DIR_SUFFIX
channel installs and profiles/-rooted HERMES_HOME values. Both read
config.yaml from a different home than main.ts does, silently dropping
desktop.electron_flags on the relaunch. Use the shared resolver and cover
both shapes with bundle-entry regression tests.
2026-09-27 06:26:35 -05:00
Brooklyn Nicholson
5c04455250 refactor(desktop): drop the redundant WSL check from the ozone relaunch
WSLg always sets WAYLAND_DISPLAY, which the native Wayland check already
covers, so the isWsl parameter no longer changes the result.
2026-09-27 06:26:35 -05:00
Brooklyn Nicholson
c0a0f097a8 fix(desktop): ship entry.ts as the Electron main bundle entry
The bundler rewrite predates #113247, and merging it back kept main.ts as
the esbuild entry. entry.ts, which picks the ozone platform and relaunches
before main loads, was no longer bundled: the WSLg Wayland relaunch never
ran. Bundle entry.ts again and run the built electron-main.mjs in a test
to prove the relaunch ships.
2026-09-27 06:26:35 -05:00
brooklyn!
9e7239acfd fix(desktop): pass wayland ozone on native Wayland sessions
Native Linux Wayland stayed on XWayland because the relaunch only ran for
WSLg. Append --ozone-platform=wayland when the user did not already choose
a platform. An explicit x11 hint and desktop.electron_flags still win.
2026-09-27 06:26:35 -05:00
teknium1
41755854ad test(e2e/windows): accept Hermes-Setup.exe's legacy receipt as a permanent shape
#125053 fixed the exe's receipt writer, but the installer .exe is not being
rebuilt, so every Desktop-installer machine keeps the old receipt
(completedAtUnix, and a null pinnedCommit without git on PATH) until
`hermes update` rewrites it. No reader depends on completedAt or
pinnedCommit: Desktop's launch gate is runtime usability, and every other
reader only checks that the file exists. So the stamp phase stops
presenting this as a gate that lifts when a fixed exe ships and asserts the
contract that holds with the released exe: only those two FAIL lines, an
integer completedAtUnix, and the checkout at the expected commit. Any
other FAIL line stays red.

Refs #124949, #125053.
2026-09-27 04:15:09 -07:00
teknium1
61b97ffc12 test(e2e/windows): gate the Hermes-Setup.exe receipt on #124949
With one PM store per leg the Desktop app no longer re-bootstraps on first
launch, so the receipt Hermes-Setup.exe writes survives to verify-stamp:
pinnedCommit null on a git-less machine and completedAtUnix instead of
completedAt, replacing install.ps1's correct receipt. Gate only that
writer's exact two FAIL lines, and then require the checkout itself to be
at the expected commit. Any other FAIL line stays red.
2026-09-27 04:15:09 -07:00
teknium1
85a1dd1cd2 test(e2e/windows): derive HERMES_HOME from a normalized workroot
With one store per leg, HEAD's install.ps1 never finished on Windows: the
source completion relaunched its store Python into itself every ~19 s, 35
levels deep after 12 minutes (job 108573522351). The workflow's workroot is
<workspace>\..\hermes-desktop-gui-e2e, so the store path carried the "..";
Windows reports sys.executable normalized, and prepare_launch()'s
Path.absolute() comparison never matched (#122513, fix PR #122518). The
job's own store (setup-pm) had no "..", which hid this until now. Normalize
the workroot, as a user's HERMES_HOME is.
2026-09-27 04:15:09 -07:00
teknium1
3859e5c5b2 test(e2e/windows): a stuck install.ps1 fails with its process table and Python stacks
install.ps1 now runs under the same watchdog as `hermes update` (40 min),
and the watchdog adds a py-spy stack for every Python process of the leg
before it stops them. The workflow installs py-spy beside pywinpty.
2026-09-27 04:15:09 -07:00
teknium1
93105d7d87 test(e2e/windows): undo installer CRLF churn through a pathspec file
`windows: {installer-script, installer-script+desktop, desktop-installer@latest}
-> hermes-desktop-app-update (v2026.7.1 -> HEAD)` failed with "Program
'git.exe' failed to run: The filename or extension is too long". A v2026.7.1
clone leaves hundreds of CRLF-churned paths, and Clear-HistoricalInstallerChurn
passed every one to `git checkout --` on the command line. Pass them through
a NUL-separated --pathspec-from-file with --literal-pathspecs instead.
2026-09-27 04:15:09 -07:00
teknium1
840f0fc9cc test(e2e/windows): one PM tool store per leg
`windows: desktop-installer@latest -> hermes-update (HEAD -> NEXT)` and
`installer-script+desktop -> hermes-update (HEAD -> NEXT)` failed with
"source launcher publication failed". install.ps1 and `hermes update` ran
with setup-pm's HERMES_RUNTIME_DIR, while the desktop smoke launched the app
without it (as a user's app runs) and settled the checkout onto a second
store at <HERMES_HOME>\tools. The update then re-pointed its own running,
locked hermes.exe.

windows-e2e.ps1 now drops HERMES_RUNTIME_DIR, HERMES_PYTHON and VIRTUAL_ENV
on entry, so every product step of a leg (installer, app, update, chat)
resolves the one store a user's machine has. The driver keeps setup-pm only
through $DriverPython (and PATH, which older installers need for uv and
ripgrep).

`hermes update` also runs under a 45-minute watchdog. Past it, the driver
writes every process (pid, parent, start time, command line) to
logs\update-hang-processes.txt, stops this leg's processes and fails with
that table, instead of being cancelled blind at the job cap.
2026-09-27 04:15:09 -07:00
teknium1
08f0691fa3 Revert "test(e2e/windows): hermes update follows the store the desktop smoke settled on"
This reverts commit c7883282c91. On real Windows (runs 36293787001 and
36293789061, c7883282 + #124702), dropping HERMES_RUNTIME_DIR for the update
turned a 40-second red ("source launcher publication failed") into a silent
40-minute hang: the update finished its code, Desktop rebuild and skills
steps, printed the curator notice, and never exited until the job was
cancelled. The store it then used (<HERMES_HOME>\tools) was created by the
desktop smoke and never received the install's PM tools, so the update's
post-update default-tool install ran against a half-populated second store.
Both one-sided variants (keep the variable in the smoke; drop it for the
update) fail. The real fix is one store per leg, and it touches every
Windows leg, so it goes in on its own with a full Windows run.
2026-09-27 04:15:09 -07:00
teknium1
bb1e5680d6 ci(install-e2e): give Windows legs a 90-minute job budget
windows: desktop-installer@latest -> hermes-update (v2026.9.24 -> HEAD)
was cancelled at the 60-minute cap (job 108542725043) with nothing wrong:
its hermes update ran the historical venv->PM takeover, npm ci and a
Desktop rebuild in 24 minutes (one npm ci stalled 11 of them), exited 0 and
passed every post-update check, and the desktop smoke was cut off. Green
Windows legs already take up to ~40 minutes and the app-update driver may
wait 30. Real hangs stay bounded by pty-run.py and the per-step budgets.
2026-09-27 04:15:09 -07:00
teknium1
2d6620f668 test(e2e/windows): hermes update follows the store the desktop smoke settled on
`windows: desktop-installer@latest -> hermes-update (HEAD -> NEXT)` and
`installer-script+desktop -> hermes-update (HEAD -> NEXT)` failed with
"source launcher publication failed". The desktop smoke runs the app
without the job's HERMES_RUNTIME_DIR, the way a user's app runs, so it
settles the checkout onto <HERMES_HOME>\tools and republishes .hermes\bin
for that store's Python. The driver's `hermes update` then ran from that
hermes.exe with the job's store (setup-pm\tools) back in scope. The update
re-pointed both launchers, and Windows refused to replace the running
hermes.exe.

Invoke-HermesUpdate now drops HERMES_RUNTIME_DIR while <HERMES_HOME>\tools
exists, so the update stays on the one store the install actually uses.
Keeping HERMES_RUNTIME_DIR in the smoke instead fixed these two legs but
broke eight others whose app never reached its backend (run 36288957697),
so that variant was reverted.
2026-09-27 04:15:09 -07:00
teknium1
a965c4052e test(e2e/windows): the app-update driver waits out the real update and checks staged main
Once launch capture works, two Windows hermes-desktop-app-update failures
show up (run 36286580917):

- v2026.9.24 -> HEAD (installer-script, installer-script+desktop,
  desktop-installer@latest): the app's update is correct but slow. From
  v2026.9.24 it runs the historical venv->PM takeover and a full Desktop
  rebuild (the `hermes update` alone took 9m43s). launch-from-spec's 10 min
  default gave up 9m55s after the Update now click while the updater window
  showed "Updating code and dependencies 9m 22s elapsed". windows-e2e.ps1
  now passes --timeout-ms 1800000, in line with open-app-update's 35 min
  wait. The driver's self-deadline is now measured after the update wait
  instead of being a flat 20 min inside it.
- HEAD -> NEXT: Desktop offered real GitHub main (0f4a98f8) instead of the
  staged NEXT, so assertStagedBranch refused to click. HEAD's Electron main
  is one ESM bundle, and checkout-source.ts binds promisify(execFile) at
  load, before installSourceBranchProbe patches execFile in the packaged
  app. POSIX avoids this by wrapping the launcher script, which Windows
  skips. The probe now also hooks ChildProcess.prototype.spawn, which every
  child passes through, and reuses branchProbeArgs, which already knows both
  the --run-module and cmd.exe .cmd shapes. Verified locally: a promisify
  captured before the hook now gets `--git <real> --branch main`, and other
  commands are untouched.
2026-09-27 04:15:09 -07:00
teknium1
03ffc307ee ci(install-e2e): a rate-limited result chart no longer turns the run red
The Result chart job lists the run's jobs and artifacts through the repo's
GITHUB_TOKEN, which shares one hourly budget with every other workflow. In
a busy hour it answered "API rate limit exceeded for installation" and
failed a run whose legs were all green (36283843710, 36285499760). Retry
after 30/60/90 s. If the API still refuses, write a warning and a summary
line instead of failing. Each leg's own conclusion still stands.
2026-09-27 04:15:09 -07:00
teknium1
9f78424dbb test(install-e2e/windows): capture the packaged launch of a PM-launched hermes desktop
`windows: * -> hermes-desktop-app-update (HEAD -> NEXT)` failed with
"a launch was actually captured (exit 0 without a launch must not pass)":
`hermes desktop` built the app and launched the real Hermes.exe, and the
driver's hook never saw it. There were two gaps, and either one alone is
enough to miss the launch:

- HEAD installs publish the PM launcher (.hermes\bin\hermes.exe), which
  runs its interpreter with -I and drops PYTHONPATH, so sitecustomize never
  loads. windows-e2e.ps1 now does what installer-script-e2e.sh already does:
  run launch-capture/pm-launch.py for a PM launcher, and keep PYTHONPATH
  for pre-PM venv console scripts.
- Since 3ddf82fe24 the Windows packaged launch is a detached
  subprocess.Popen, not subprocess.run. sitecustomize now also intercepts a
  launch-shaped Popen: it records the same spec and stands in for a child
  that already exited with 0. Every other Popen, including the one inside
  subprocess.run, spawns for real.
2026-09-27 04:15:09 -07:00
teknium1
5f9c5a2c38 test(install-e2e/windows): tolerate v2026.9.24's CRLF clone churn before the GUI update
Three `-> hermes-desktop-app-update (v2026.9.24 -> HEAD)` legs failed with
"installed source has only installer-generated changes (other: 1 .M N...
<same blob> <same blob> website/i18n/zh-Hans/.../user-stories.mdx)".

v2026.9.24's install.ps1 clones under Git for Windows' system
core.autocrlf=true and pins false only afterwards, so the ~1700 text files
.gitattributes does not pin to LF (*.mdx, LICENSE, ...) sit on disk as CRLF
over LF blobs. The stat cache hides them, except the last few the clone
wrote inside the index's racy timestamp window, which read as modified.
Which files show up varies (1 or 3 per leg); main's installer no longer
pins autocrlf.

Clear-HistoricalInstallerChurn now also restores a `1 .M` entry whose
index and worktree modes match and that has no `diff --numstat
--ignore-cr-at-eol` record (a line-ending-only difference). A real content
edit, an edit inside a CRLF file, a mode change, or an untracked file still
fails the assertion, listed.
2026-09-27 04:15:09 -07:00
teknium1
5afba1e8d9 ci(install-e2e): keep one install-e2e-red issue in step with the scheduled matrix
A red scheduled run blocked nothing and told nobody. install-e2e-red.yml
runs after every scheduled "Install & Update E2E" run: red opens one issue
labelled install-e2e-red (or rewrites the open one's body in place) with the
red legs grouped by failure class and linked; the first green run closes
it. No per-run comment, never a second issue.

It is its own workflow_run workflow because install-e2e.yml is also called
by stable-release.yml with read-only permissions, and a nested job asking
for issues: write would fail that call at startup. workflow_dispatch with a
dry-run default previews the change for any run id.
2026-09-27 04:15:09 -07:00
teknium1
7605349f25 ci(install-e2e): run a path-filtered four-leg subset on pull requests
install-e2e.yml only ran on the clock, so nothing in front of a merge
installed a release and updated it on a real OS. A pull_request trigger,
path-filtered to the install/update surface (derived from 60 days of
update/install/pm commits), runs the new `pr` route of
generate-e2e-matrix.mjs with only the newest release tag sampled:

  linux   installer-script -> hermes-update (newest release -> PR)
  linux   installer-script -> hermes-update (PR -> NEXT)
  windows installer-script -> hermes-update (PR -> NEXT)
  macos   installer-script -> hermes-update (newest release -> PR)

The bundle-manifest validation job is skipped on PRs (bundled legs need
dispatch-only manifests). The full matrix stays on schedule and release.
2026-09-27 04:15:09 -07:00
teknium1
116b3e2f9e test(install-e2e/windows): the first chat turn exits under the pseudoconsole
Every Windows leg of every scheduled run since 2026-09-25 was cancelled at
the 60-minute job timeout in its Install step. The user-state phase runs
`hermes chat -q "..."` under pty-run.py (a ConPTY, so the CLI sees a real
TTY). Since a5c7eed (shipped in v2026.9.21) `-q` on a TTY seeds an
interactive session and only answers-and-exits with --oneshot; the turn
answered ("Hello from the mock inference server!") and then sat at the
prompt. pty-run.py's --timeout 900 never fired either: it checked the
deadline only between blocking PtyProcess.read() calls, and an idle prompt
never returns from read().

- windows-e2e.ps1 passes --oneshot when `chat --help` advertises it (older
  tags have no flag and exit after -q on their own), lowers the turn budget
  to 300s, and names a 124 as "never exited" instead of a generic failure.
- pty-run.py drains the pty on a daemon thread and waits on the queue with
  the deadline, so --timeout holds whatever the child does, and writes a
  TIMEOUT line into the captured log before killing it.
2026-09-27 04:15:09 -07:00
teknium1
d25bbd01b7 test(e2e/network): drop the #124653 gate; the channel 503 cell passes
The known_failure wrapper on test_channel_503_with_retry_after_is_retried_then_updates keyed on the one-attempt channel read. With update-path retries it passes, so the cell asserts plainly again.
2026-09-27 03:41:42 -07:00
teknium1
d1484ab44b fix(update): only an explicit hermes update retries channel reads
Passive checks (`hermes --version`, the banner, the Desktop/dashboard
update check) now make one channel-read attempt again. Offline usually
surfaces as DNS EAI_AGAIN or ENETUNREACH, which pm.network.is_transient
classes as transient, so wrapping every read in retry_network added 7 s
of backoff to the synchronous version line, stretched a hung CDN from
30 s to ~127 s, and logged a WARNING to stderr on every retry.

`release_channels.retrying_reads()` opts a block in; update_cmd wraps
the channel resolution of `hermes update` in it. Tests: keep the
Retry-After retry test (now scoped to the update path) and replace the
budget-exhaustion test with one pinning that passive reads and 404s make
exactly one attempt.
2026-09-27 03:41:42 -07:00
JoaoMarcos44
0da04c192c fix(update): retry transient release channel reads 2026-09-27 03:41:42 -07:00
teknium1
54388a21ca fix(cli): compare interpreter paths lexically, not through symlinks
resolve() on both sides also collapses a venv interpreter onto the store
binary it links to, so a caller still running in its old venv stopped
re-executing into the store interpreter (two tests in
tests/pm/test_source_update_launch.py go red). The loop in #122513 is a
spelling difference, not a symlink: PM spells the store Python through a
HERMES_HOME that may contain '..', and the OS reports sys.executable
normalized. normcase(abspath()) matches those spellings and keeps a venv
interpreter distinct.

Tests: the '..' spelling is current (red on main); a venv python symlinked
to the store binary still relaunches (red with resolve()).
2026-09-27 03:38:52 -07:00
Konstantin Khlopkov
1d6e53eb05 fix(cli): compare resolved interpreter paths so a current venv launch does not relaunch 2026-09-27 03:38:52 -07:00
teknium1
5a0225dfff fix(gateway): restart watcher names its checkout instead of inheriting the cwd
The detached watcher runs as `<python> -c <program>`, so `hermes_cli` resolved only
because update_completion happened to spawn it with cwd=<checkout>. The program now puts
the checkout on sys.path itself, and the bare-Python test runs it from an unrelated cwd
(red without the sys.path line).

Also drops the two unused re-export aliases in gateway.status (`_posix_is_zombie`,
`_pid_exists_win32_ctypes`): nothing imports them and neither is in the old-updater
compat surface.
2026-09-27 03:37:09 -07:00
teknium1
029445545c fix(gateway): keep the update restart watcher stdlib-only so it survives the bare store Python
After the package-manager handoff, hermes update finishes on the bare store
interpreter and spawns the detached restart watcher as sys.executable -c.
The watcher imported gateway.status (utils -> hermes_yaml -> ruamel) and
hermes_cli.config, died with ModuleNotFoundError before relaunching, and
left every manually started gateway down after the update (gated on #124649).

Move the stdlib liveness probe (zombie-aware POSIX kill(0), Windows
OpenProcess) into hermes_cli._subprocess_compat, have gateway.status's
fallback delegate to it, and import only stdlib-backed modules in the
watcher.

Fixes #124649.
2026-09-27 03:37:09 -07:00
teknium1
07be7d47ed fix(bootstrap-installer): write the shared receipt schema and resolve the commit without git
Hermes-Setup.exe replaced install.ps1's .hermes-bootstrap-complete with its
own copy: completedAtUnix instead of the completedAt every other writer and
reader uses (install.ps1, Electron main, hermes_cli/source_stamp.py,
scripts/verify-bootstrap-version-stamp.py), and a pinnedCommit from a bare
`git rev-parse HEAD`, which is null on a machine where git exists only as
the PM-staged binary install.ps1 uses.

The receipt now carries completedAt (UTC ISO-8601, milliseconds), and the
commit resolves like install.ps1's Stage-Complete: the pinned commit, else
the checkout HEAD read from its ref files (loose ref or packed-refs), else
the full sha install.ps1 already recorded (read BOM-tolerant).

Fixes #124949
2026-09-27 03:35:26 -07:00
teknium1
80512c90dc fix(website): batch the plugin stars probe at 100 repos per GraphQL request
A single request for all 317 catalog repos now exceeds GitHub's per-query
resource limit: the reply carries partial data plus an error, the script
prints 'Probed 317', and every repo past the limit keeps its stale count
(hindsight showed 26.5k stars while GitHub had 34.5k). Batch at 100; probed
live: 317/317 repos in 4 requests, no errors.
2026-09-27 03:34:56 -07:00
teknium1
e2abb460b6 fix(website): put the repo root on sys.path before importing hermes_yaml in the stars probe
The scheduled Build Skills Index run executes
`python website/scripts/fetch-plugin-stars.py --probe`, so sys.path[0] is
website/scripts and the repo-root `hermes_yaml` shim it now imports is not
importable. Every scheduled probe since 2026-09-24 died with
ModuleNotFoundError, the live plugin-stars.json froze at that date, and the
41 catalog entries added since have no star count: they sort to the bottom of
'Most starred' and show no star pill regardless of real stars.

extract-plugins.py already inserts REPO_ROOT; do the same here.
2026-09-27 03:34:56 -07:00
teknium1
afc05274a2 website: keep the plugin name whole; tier/version pills move to the chip row, tools as one ellipsized list
With Community + version + stars pinned beside the title at 3 columns the name
collapsed to 'hermes-r…' — the one element a card must never lose. The title
now owns its row (stars stay right-aligned); tier and version pills lead the
chip row. The Tools row renders the first three names as one comma-separated
run that ellipsizes as a whole plus '+N', instead of three shrinking chips that
read 'mnemosyn… mnemosyn… mnemosyn…' for same-prefix tools.
2026-09-27 03:34:51 -07:00
teknium1
97a3607838 feat(website): uniform plugin catalog cards (clamped text, capped tool list)
Every card in a Plugin Catalog grid now renders at one height, so rows line
up instead of staggering by description length, banner presence and tool
count (the Desktop section had 28 distinct card heights, 412px to 1485px).

- grid-auto-rows: 1fr + column-flex card; the detail block (facts, install
  box, links) sinks to the bottom edge of every card
- banner slot is always 120px: entries without an image get a placeholder
  (tier tint gradient + large muted category glyph); a broken image URL
  swaps to the placeholder instead of collapsing the slot
- title row is one line (ellipsis), description clamps to 3 lines with an
  ellipsis and reserves 3 lines when shorter; full text stays in `title`
- chip row and dates row are one fixed line each (right-edge fade mask on
  chips); dates row renders empty when undated
- facts block is exactly three fixed-height rows: Maintainer, Pinned (with
  `Requires` folded in as `· hermes >=x`), Tools — first 3 tool names as
  chips plus a dashed `+N` chip (full list in `title`), or a muted dash
- install command truncates with an ellipsis instead of scrolling

`capToolChips` in PluginCatalog/catalog.ts is the cap helper (first 3 + N).
Desktop picker embed (`?embed=picker`) uses the same card and stays uniform.
2026-09-27 03:34:51 -07:00
teknium1
a3f454a287 fix(build): keep scripts/build/inputs.py free of pm imports; accept musl targets in its grammar
The Nix agent derivation builds scripts/build/*.py from a fileset that
does not include pm/, so importing pm.store.ALL_TARGETS there failed
nix flake check with ModuleNotFoundError. Extend the local target
regex with linux-(x64|arm64)-musl instead.
2026-09-27 03:24:39 -07:00
teknium1
aba7deac48 test(e2e): musl host cell passes; drop its #123682 known_failure gate
On Alpine the installer now stages musl uv/Python/Node (a full install
completes once libstdc++ is present) and, on the stock bash/git/curl
image, refuses up front naming musl and libstdc++.
2026-09-27 03:24:39 -07:00
teknium1
842f162f7d fix(installer): refuse musl hosts without libstdc++ up front, naming it
PM's musl Node is the unofficial-builds musl archive, which links the
system libstdc++. On a stock Alpine (bash, git, curl) node and npm then
fail staged verification with raw relocation errors after the clone and
downloads. Check in the prerequisites stage and name the package.
2026-09-27 03:24:39 -07:00
teknium1
5cbca02510 test(pm): keep FFmpeg's musl gaps when the re-pin test narrows its gap table
The test replaced FFmpeg's gaps with only the bionic entry, which now
turns the musl targets back on and asks BtbN for a glibc build that
FFmpeg.fetch_url refuses on musl. Extend the real table instead.
2026-09-27 03:24:39 -07:00
teknium1
f5c3c7d70a test: trim musl coverage to two invariants
Keep one PM invariant (a musl ELF userland resolves to linux-<arch>-musl
and the default closure is satisfiable with musl-native artifacts) and
one installer invariant (uv_bootstrap_target picks musl uv from a musl
ELF interpreter even when ldd says glibc). Drop the lock-content and
URL-shape change-detectors.
2026-09-27 03:24:39 -07:00
teknium1
34d52883cb fix(installer): ELF interpreter decides musl in uv_bootstrap_target, as in pm/store
install.sh consulted ldd before the ELF interpreter while pm/store.py's
_is_musl_libc reads the native userland's ELF interpreter first, so a
glibc ldd (secondary toolchain, gcompat) could make the bootstrap stage
a different libc than PM later resolves. Read /bin/sh (then /bin/ls)
PT_INTERP first in both; ldd and the loader glob are fallbacks only.
2026-09-27 03:24:39 -07:00
teknium1
ee0ad5adb2 refactor(pm): expose MUSL_TARGETS publicly; drop the per-URL pin hash cache
packages.py and security_packages.py imported the private _MUSL_TARGETS
across modules; make it a public store constant next to ALL_TARGETS.
The _pin_tool hash memo is an unrelated optimisation, out of scope here.
2026-09-27 03:24:39 -07:00