Commit Graph

2576 Commits

Author SHA1 Message Date
ethernet
e336fbcf79 fix: release cut output waits for its CI run and leads with the next action
The cut printed before GitHub listed the dispatched run, so it almost always
said "the run is not listed yet". It now polls up to 30s for the run, and when
none lists it links the stable-release workflow page. The copy is shorter:
Release CI, the skipped-tests warning (bold on a tty), the draft link, and the
publish command on its own line.
2026-09-29 15:12:04 -04:00
ethernet8023
792ba0a264 fix(release): read a tag's record, not its signature
git keeps an annotated tag's signature inside the message body, so
`git tag -l --format=%(contents)` returns the JSON record followed by an
armored block, and `json.loads` of the whole message raises "Extra data: line 2
column 1" at the armor's first byte. Claim tags, final receipts and build
receipts are created on a host whose git signs tags, so stable admission
(`python -m scripts.releases.stable admit`), ordered reconciliation
(`python -m scripts.releases.sequencer`) and tagless build receipts each
refused to read their own tag.

Add `versioning.tag_record()` as the one reader of a tag's record — the text
before the armor, unchanged for an unsigned tag or a detached signature — and
route every `%(contents)` read through it in stable.py, sequencer.py and
commit_build.py. commit_build imports it inside the function so its module
surface stays stdlib-only for the isolated checkout-admission step.

The reads inside the release tests use the same reader, and the one fixture
that creates a plain final tag says so explicitly, so the release suite passes
on a host whose git signs tags by default.
2026-09-29 14:55:15 -04:00
ethernet
36699ac97a Merge pull request #128156 from NousResearch/fix/desktop-flavored-icons-clean-checkout
Desktop release builds stop leaving their flavored icon render in the checkout
2026-09-29 13:08:09 -04:00
ethernet8023
4bc38bceec fix(desktop): hand the admitted checkout back after flavored icon staging
electron-builder resolves every piece of packaging artwork from the workspace
— the app icon, the icon.ico extraResource, the MSIX logo set before-build
stages into build/appx, and the exe identity stamp — so a flavored canary or
commit build copies its rendered icon set over apps/desktop/assets before
packaging.

Those files stopped being gitignored in 4b7229d612, which turned a write that
was invisible into 15 dirty paths. The source custody check then refuses the
checkout twice: inside the same build, at the check that used to sit after the
copy, and on every later build from the first check on a now permanently dirty
tree.

Hold the render at the workspace path for the packaging window only and
restore the admitted files afterwards, whether packaging finished or failed.
The final custody check now asserts the build handed the checkout back.
2026-09-29 12:29:47 -04:00
Mark S.
ca705dbf7e fix(ci): card-tool contract test reads the as-const list and runs on desktop edits (#128010)
#127828 changed CARD_TOOL_NAMES in apps/desktop/src/lib/tool-render-class.ts
from `new Set([...])` to `[...] as const`. The pytest that pins that list
against the gateway's _TOOL_LIFECYCLE_UI_TOOLS located it with a regex that
only matched `new Set([`, so it has failed on main since with
"CARD_TOOL_NAMES not found in tool-render-class.ts". The gateway set still
covers every desktop card name; only the test's parser was stale.

The PR merged green because it touched only apps/, which skips the Python
lane, so the one test that reads that file never ran on it.

- The parser accepts both declaration forms and fails on an empty list
  instead of passing a vacuous subset check.
- tool-render-class.ts joins _PY_RELEVANT_CONTRACT_FILES so an edit to it
  runs the Python lane, with a classifier case for the entry.
2026-09-29 12:27:26 -04:00
ethernet
74163fa18c fix(desktop): pass release epoch to native version subprocess 2026-09-29 10:55:18 -04:00
kshitijk4poor
596678d665 refactor(update): tighten the partial-clone heal after review
- _deepen_shallow_repo marks packs in its finally without a git config probe
  that could itself raise and turn a finished unshallow into False; markers
  are inert in a clone without a promisor remote.
- Installers warn instead of aborting the stage when a marker cannot be
  written, matching the Python helper.
- Docs: the manual commands follow HERMES_HOME, and the missing-object
  backfill feeds IDs through `fetch --stdin` so a large list cannot exceed
  the Windows command-line limit.
- Test and docstring comments describe the contract, not the incident.
2026-09-29 15:24:15 +05:30
kshitijk4poor
7b20b72269 fix(install): installer rerun heals a partial clone stuck on the pack-objects crash
An install whose updater dies on the git 2.53+ pack-objects BUG (#124272)
never fetches the release that heals it: the fetch runs in the installed
code. The installer rerun (install.sh / install.ps1, served fresh) is the
path that runs new code, so when origin is a promisor remote it now marks
the checkout's unmarked packs before its existing-checkout fetch. Marking is
idempotent and never rewrites objects, so it runs up front rather than as a
retry that would first print a failed-fetch report.
2026-09-29 15:24:15 +05:30
teknium1
9b7db3105b fix(telemetry): fewer rows per busy day and a smaller local store
A volume simulation through the real store showed two costs worth cutting.

Rows: hermes.task_run.finished carried duration, retry, model-call and
tool-call buckets on top of outcome/end reason/surface, so almost every task
became its own row; hermes.tool_call.count did the same with latency and
retry buckets. The terminal rows now keep only the fields they are filtered
by, and the same end events feed two small split counters:
hermes.task_run.duration (surface, outcome, duration, retries) and
hermes.tool_call.latency (tool category, latency, retries). Per-task call
counts already ride on hermes.task_cost.count. v3 still accepts the v2 field
sets of both counters so rows recorded before an upgrade drain.

Local store: every package lived three times for 30 days (counter rows, the
payload column, a pretty-printed outbox file). Outbox files are now compact
JSON, and once the ingest has accepted or refused a package the database
drops its copy of the body; the file stays as the local history the docs
promise, and pending packages keep their body for resends.

Measured (simulated day, skewed draws, before -> after):
  heavy    2,788 -> 2,347 rows, 21.0 -> 17.7 KB on the wire, 102 -> 44 MB local over 30 days
  extreme  6,561 -> 5,047 rows, 46.0 -> 33.6 KB, 272 -> 102 MB
  typical    663 ->   650 rows,  7.1 ->  6.8 KB,  24 ->  13 MB
2026-09-28 12:43:03 -07:00
teknium1
0b6ac0b622 fix(telemetry): live relay smoke no longer expects adoption from an agent-created skill
The signals fix (M15) stopped counting Hermes' own skill creation as user
adoption, so the smoke's agent-created skill records no feature_adoption
row; the exact name sets now assert its absence.
2026-09-28 12:43:03 -07:00
teknium1
d6e4e33d7d fix(telemetry): live relay smoke expects the v5 efficiency and adoption rows
The smoke's exact metric-set checks (SQLite store and export packages) were
never updated for v5: every attended CLI turn now also emits task_cost,
tool_output_truncation, tool_overhead and tool_enabled_unused, and the skill
lifecycle step latches feature_adoption(skills_created). A correct run failed
with "Unexpected SQLite counters", so the documented live gate could no
longer catch a real regression. Both sets now include the five metrics and a
shared _validate_v5_rows asserts their dims in the store and the packages
(task_cost provider/model `custom` for the canary model, api_calls 2,
tool_calls 1; read_file truncation; file toolset used; tool_overhead on cli).

Live smoke (loopback fake OpenAI server, no model spend):
  before: rc=1 AssertionError: Unexpected SQLite counters
  after:  rc=0 "Hermes -> NeMo Relay shared-metrics smoke test passed"
Sabotage: expecting the canary model id in task_cost is caught.
2026-09-28 12:43:03 -07:00
teknium1
c5555efca8 test(smoke): expect the per-response model_reply_issue counter
Every primary model response now records hermes.model_reply_issue.count
(issue=none for usable replies), so the smoke turn's two scripted replies
add one row with value 2; the exact-metric-set assertions must include it.
2026-09-28 12:43:03 -07:00
teknium1
b1e9c2a5a1 test(smoke): expect v4 tool-quality, context-peak and startup counters 2026-09-28 12:43:03 -07:00
teknium1
ccd1a274e8 fix(telemetry): user-named providers and models never leave the machine; milestones use the owning profile
Independent review of the v3 metrics found:

- Provider fields were shape-checked only, so `custom:<config key>` (and
  any unshipped provider id) reached setup.completed, install.snapshot,
  model_switch, fallback, model_tokens and model_route. Providers now pass
  only when Hermes ships them (auth registry, overlays, model catalog,
  aliases, cached models.dev ids); everything else reads `custom`. A custom
  or loopback provider's model id reads `custom`, as does any model id that
  looks like a path or URL. Two existing tests asserted the old export of a
  custom endpoint's model id and an unshipped provider; they now assert the
  collapse, and the smoke run treats the model canary as prohibited.
- Install milestones read the install age on the Relay thread, which has no
  profile binding, so a multiplexed profile got the launch profile's age.
  The subscriber captures its home at construction.
- Gateway sessions retired by the store (auto-reset, /new, resume) now close
  their metrics session at the route transition instead of at shutdown.
- The dashboard MCP add records its install after leaving the config lock.
- Compression and gateway slash counting helpers move into
  shared_metrics_events so the >2k-line files barely grow.
2026-09-28 12:43:03 -07:00
teknium1
3d62ae2233 feat(telemetry): decision-data docs, consent copy, smoke coverage and install-age for new installs
- Docs: a Decision-data metrics section mapping each metric to the product question it answers;
  skill-load and snapshot paragraphs describe the new public-name and install-age fields.
- setup consent copy names the new data classes (session length, token totals, command and
  catalog names, bucketed setup counts).
- The real-CLI smoke now asserts the session summary, token sums, TTFT bucket and milestones in
  both the SQLite store and the exported, schema-validated package.
- A profile with no sessions yet is a brand-new install (lt_1h), not unknown: the first task's
  milestone and the first snapshot fire before state.db has a session row.
- Two invariant tests: session summaries + once-per-install milestones; token sums per model and
  auxiliary task with private task names collapsed.
2026-09-28 12:43:03 -07:00
teknium1
b8732b768e feat(telemetry): shared metrics v3 with failure classes, tool usage, platform and daily snapshot
The fleet rollups could say that tasks and tools failed, but not why, on which
model, on which messaging platform, or which features installs actually use.

- model_route rows gain call_role, outcome and error_class (the classifier's
  FailoverReason for the last failed attempt; a success row keeps the error it
  recovered from).
- task_run rows gain platform (built-in gateway platform, plugin, or none);
  task_run.finished gains failure_class from a closed set.
- new hermes.tool.usage.count: tool_name (only names from the static
  toolsets.TOOLSETS snapshot; mcp/plugin otherwise), outcome, error_class.
- new hermes.install.snapshot, latched to once per rolling 24h like
  client.active: memory provider plus bucketed counts of MCP servers,
  plugins, skills, cron jobs and profiles; no names.
- package schema v3; v2-shaped pending counters still validate and drain.
- consent copy and developer docs list the new categories; the relay smoke
  asserts v3 shapes and disables title generation (its third model request
  broke the smoke's request-count check on main).
2026-09-28 12:43:03 -07:00
Teknium
a5bd246865 Old pre-decomposition import paths are gone: plugin compat layer removed on schedule (#126164)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule

The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:

- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
  names, restored dead definitions) and the three re-export stub modules
  (gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
  the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
  `plugins.allow_deprecated_imports` escape hatch

An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.

hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).

In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.

* chore: retrigger CI (zero-job startup_failure phantom)

* test: drop resolution allowlist rows for the two deleted which() sites

hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
2026-09-28 10:21:41 -07:00
alt-glitch
1e2498a23d fix(whatsapp): stamp every bridge console line in the local log shape
The adapter captures the bridge's stdout and stderr verbatim into
bridge.log, so a console line without a stamp cannot be placed in time.
The previous commit stamped 12 of the bridge's console calls one by one
with a bracketed UTC ISO stamp; the other calls (config summary, pairing
mode, reconnect-scheduler and version-resolver messages, inbound media
download failures) stayed bare.

This replaces the per-call wrapping with one installConsoleStamps() call
at bridge startup that prefixes every console.log/warn/error line with
`YYYY-MM-DD HH:MM:SS,mmm ` in local time, the shape the Python logs in
logs/ already use. The 12 call sites go back to plain console calls.

Lines that programs parse must stay bare. The dashboard pairing watcher
runs json.loads on each stdout line of `bridge.js --pair-json`, so the
JSON event lines (pair events, debug, poll_update_decode and `ignored`
events) now go through writeJsonLine(), which writes to process.stdout
directly. The terminal QR block is written the same way so its first
row is not shifted. Pino (the Baileys logger) writes to fd 1 itself and
never passes through console, so its lines are unchanged.

Refs #97021
2026-09-28 18:21:25 +05:30
ygd58
ee2c8b52a6 fix(whatsapp): timestamp bridge.js's lifecycle console.log/warn lines
Fixes #97021.

The WhatsApp bridge's human-facing lifecycle lines -- startup
("bridge listening", "session stored"), connection state (connected,
logged out, restart/reconnect), and "[bridge] ..." warning lines --
were bare console.log/console.warn output with no timestamp. The
platform adapter pipes the bridge's stdout/stderr verbatim into
bridge.log, so no timestamp is added downstream either. Only
structured JSON events (pair events, allowlist rejections from #92683)
carried a timestamp. During incident forensics this made bridge.log
impossible to sequence on its own -- the exact durations of a
relaunch/re-pair loop, or the ordering of "Logged out" relative to a
fatal stream error, couldn't be reconstructed without cross-correlating
gateway logs and the sparser pino output.

Added timestampedLine() to bridge_helpers.js (the existing pure,
testable helper module bridge.js already draws from for reconnect
scheduling and version resolution) -- a small function that prefixes a
message with an ISO-8601 UTC timestamp in brackets, matching the
issue's own suggested fix direction and kept deliberately distinct
from the `ts: Date.now()` epoch-ms convention #92683 already
established for the JSON event stream (these are plain human-readable
lines, not JSON payloads).

Applied it at the exact lifecycle line sites the issue's own line
inventory cites: the core startup lines ("bridge listening on port",
"session stored in"), the connection lifecycle lines (connected,
logged out, restart-code-515, reconnect-in-3s, pairing complete), and
the five "[bridge] ..." warning lines (poll update/upsert aggregation
failures, gif/ffmpeg conversion fallback, failed read receipt).
Scoped narrowly to what the issue asked for -- did not touch the
already-structured JSON event lines, or the separate config-summary
startup block (allowed users / DM policy lines) the issue didn't cite.

Added a new test file (bridge_helpers.timestamp.test.mjs), following
the established plain-assert, no-framework pattern from the existing
bridge.reconnect.test.mjs: verifies the prefix is a valid, parseable
ISO-8601 timestamp; the original message text (including
template-literal-interpolated content, matching the actual call sites)
survives byte-for-byte after the prefix; and two calls a moment apart
produce non-decreasing timestamps, confirming the prefix reflects
actual call time rather than a cached value -- so a relaunch loop's
individual lines stay independently sequenceable, the issue's core ask.

All assertions pass in the new test file. Ran the existing
bridge.reconnect.test.mjs, bridge.sendqueue.test.mjs, and
allowlist.test.mjs directly -- all pass unchanged (no regression to
bridge_helpers.js's other exports or to bridge.js's own logic, since
this only wraps pre-existing message strings passed to console.log/
console.warn without changing any control flow).
2026-09-28 18:21:25 +05:30
teknium1
8a3ede1be0 fix(bundle): payload smoke runs the harness with telemetry off, like runtime
browser-harness sends each CLI event from a detached python child; in the
smoke it outlived the check and held the relocated payload open on Windows,
so the restore rename failed with EPERM. Hermes's own harness env sets
ANONYMIZED_TELEMETRY=false; the smoke now does the same.
2026-09-27 23:53:40 -07:00
teknium1
2843d8eea4 test(bundle): smoke the payload's browser-harness on the store interpreter
The relocated-payload smoke now runs browser_exec's engine the way
tools/browser_use_cli.py launches it (store interpreter, -S, only the
payload site dir on PYTHONPATH), so a payload that cannot start the
default browser driver fails the release build instead of shipping.
2026-09-27 23:53:40 -07:00
teknium1
1d287d5375 fix(browser): ship the Browser Use CLI engine in every install, Desktop included
The default browser_exec tool ran the `browser-use` CLI from a PM side
environment (browser-use==0.13.10 in <home>/environments/browser-use),
provisioned by the installers and `hermes update`. Sealed Desktop payloads
skip that step, so the Desktop app never had it and silently fell back to
the built-in tools; the side env was also per-profile and 225 MB.

The CLI's execution path is only `browser_harness.run.main()`; the
browser-use agent framework (anthropic/openai/google-api pins, 93 MB of
googleapiclient) is never imported. browser-harness itself is 2.6 MB of
pure Python whose pins (Pillow 12.3.0, websockets 15.0.1) already match
Hermes's own, so it becomes a core dependency and runs on sys.executable:

- pyproject/uv.lock: browser-harness==0.1.13 (+ cdp-use, fetch-use).
- _find_cli() returns [sys.executable, -m, browser_harness.run]; the child
  env points PYTHONPATH at the harness site dir (the Desktop store
  interpreter boots without a venv and the harness daemon re-runs
  sys.executable), replacing whatever the agent inherited.
- The side-env provisioning (install_cli, the update/installer step) goes.
2026-09-27 23:53:40 -07:00
teknium1
b63c138d78 fix: Hermes never falls back to the user's node/npm/npx/uv
Maintainer ruling: only Hermes and only its packaged package managers
(uv/pip/node/npm) are ever used; no PATH fallback when the managed tool is
missing, no "prefer the user's if new enough".

- hermes_constants.find_node_executable: node/npm/npx resolve to PM's
  installed copy or None. Every caller already pm.ensure()s on None, so a
  missing runtime is now provisioned instead of silently borrowing the
  user's Node (native-addon ABI / npm cache mismatches).
- agent/lsp/install._install_npm: pm.ensure('npm') when PM npm is absent,
  instead of failing over to whatever npm is on PATH.
- gateway._append_node_dir_for_service: stop baking the invoker's PATH node
  dir into generated systemd/launchd units.
- main_install_repair._resolve_node_runtime_npm: drop the PATH re-scan for
  another npm.
- source_build.source_product_current: run the freshness reader only with
  PM's node.
- doctor: Node/npm rows and npm audit use PM's copies (Termux APT distro
  keeps its system Node).
- install.sh ensure_uv / install.ps1 Get-Uv: always stage the pinned uv
  artifact; delete the "uv on PATH if new enough" developer shortcut.
2026-09-27 22:04:26 -07:00
teknium1
27062c3474 fix: install cua-driver and the Browser Use CLI by default again
Computer use and browser use are meant to work out of the box. The PM rewrite
(3d12e86ef1) and the MSIX installer rework (47f4ab3a17) dropped the
install-time cua-driver fetch (7060ac7bed) and the Browser Use CLI install
(baa6b2e34d). On a fresh install the computer_use check_fn therefore stayed
False, so the tool never reached the model and its lazy ensure could not fire,
and browser_exec quietly fell back to the built-in tools.

- cua-driver is a default PM package, so the installers, a bare
  `hermes pm install` and `hermes update` carry it on every target it builds
  for. Adds the missing Android gap (the lock has no bionic artifact).
- The shared default-tool step, which the installers (via source completion)
  and `hermes update` both run, provisions the Browser Use CLI for the default
  and explicit Browser Use backends. `--skip-browser` declines it along with
  agent-browser, and `off`/Camofox never use it.
- Installers regain --skip-computer-use / -SkipComputerUse (recorded as
  `--without cua-driver`).
- The update message stops calling every default "browser tools".
2026-09-27 19:18:58 -07:00
happy5318
113a63e853 fix(plugins): make catalog known_issues informational, not a fail-closed gate
Relaxes the fail-closed install gate from #124037 per teknium1 review
2026-09-27 (known_issues informational; guard belongs at mode-selection seam):
- dashboard_install_plugin: no longer refuses known_issues entries — the text
  flows into the result warnings plus a machine-readable "known_issues" key so
  the UI can show it; the memory-provider migration paths (migrate_all_homes /
  recover_at_startup via _install_into) install hindsight again in any mode.
- cmd_install: prints the yellow "Known issue:" lines but drops the TTY
  requirement and the y/N confirmation; non-interactive installs proceed.
- entry_capability_summary: includes "Known issues: ..." so install prompts and
  the catalog UI surface the text.
- scripts/validate_plugin_catalog.py: register known_issues in KNOWN_KEYS.
- tests: keep parse round-trip + install-summary display; drop the live-catalog
  prose pin (test_live_catalog_hindsight_declares_known_issues) and the
  zoneinfo import hack.
- contributors: map 5318happy@users.noreply.github.com -> happy5318.
2026-09-27 17:06:55 -07:00
kshitijk4poor
36c91657da fix(install): kill the PortableGit extractor tree on timeout, drop dead flatten
Process.Kill() on PS 5.1 kills only the stub; its post-install children kept
%TEMP%\hermes-git-bootstrap-<PID> held past the finally cleanup. taskkill /T /F
reaps the tree. The single-wrapper-dir flatten is gone: the sha-pinned
PortableGit roots cmd\git.exe directly (verified against the real archive).
2026-09-28 02:04:08 +05:30
finn763
7a40c28750 fix(pm): execute the pinned git SFX from a scratch copy, not the cache
CI run 36189416163 failed after "git: unpacking": executing the cached
fetch-<sha> PortableGit PE in place left it handle-held (Defender
on-execute scan / the stub's RunProgram child chain) past pm's ~2 s
_remove_entry retry, so download cleanup raised WinError 32. Review
(teknium1, 5 threads) adds the rest:

- Git.unpack now copies the artifact into a .sfx-* dir beside the
  staging tree and executes the copy; pm's .staging-* teardown
  (ignore_errors) owns that path, so any hold lands on a disposable
  path and the cache dir only ever holds read handles
- refuse off-Windows with the cross-host trade-off stated instead of a
  raw PermissionError; documented in package-management.md and the PR
- the extractor is a GUI-subsystem stub that is silent under -y: error
  messages now carry the exit code and the usual causes (disk full,
  path length, antivirus) instead of promising captured output; the
  docstring no longer claims "no GUI" and records that the stub shows
  an Extracting window and runs the vendor post-install
- install.ps1: WaitForExit(600000) + Kill() mirrors pm's timeout=600,
  Fail reports the exit code and the silence
- tests: the OS-refused-exec assertion and the not-in-text change
  detectors are replaced by subprocess-argv invariants (scratch copy
  location, exit-code message, off-Windows guard before any execution)

Related to #122512

(cherry picked from commit adaf76a286bebe30c89aee4174df27f1945b60c9)
2026-09-28 02:04:08 +05:30
finn763
6d9ae2da6f fix(install): stage pinned git without tar/bzip2 (PortableGit SFX)
Windows 10 boxes whose System32 tar.exe cannot run the bzip2 filter die
at stage=prerequisites with "unable to run program bzip2 -d" while
extracting the pinned Git-2.53.0.3 tar.bz2 (#122512). Repin git for
both win32 targets to git-for-windows' PortableGit self-extracting 7z,
which carries its own extractor and the bundled usr/bin/bash.exe:

- pm/lock.json: new artifact urls + sha256 (the pin authority)
- scripts/install.ps1: generated fragment regenerated; Get-PinnedGit
  downloads the SFX and waits on it explicitly (the stub is a
  GUI-subsystem exe, so PowerShell's & does not wait); the System32
  tar invocation and its msys symlink excludes go away
- pm/packages.py: Git.fetch_url/Git.unpack run the self-extractor
  after the sha256-verified download
- pm/store.py: drop the now-callerless git_msys branch of extract_tar
- tests: RED->GREEN test runs the real Get-PinnedGit against a
  bzip2-less System32 tar.exe stub with no bzip2 on PATH; the three
  obsolete tar-contract tests and the install.ps1/PM msys-links parity
  test are replaced by a no-external-decompressor contract test

Closes #122512

(cherry picked from commit 4915304213495d3207ec6cd659e57cd16ef08d60)
2026-09-28 02:04:08 +05:30
kshitijk4poor
44e42b2502 fix(install): re-runs update tag-pinned narrow checkouts instead of aborting
The installer's re-run path fetched `origin <branch>` by name, then checked
the branch out and fast-forwarded to origin/<branch>. On a checkout an older
installer made with `--depth 1 --single-branch --branch <tag>`, the remote's
only refspec maps the tag: the fetch wrote FETCH_HEAD but no origin/<branch>,
and `git checkout main` failed with "pathspec 'main' did not match".

Fetch by explicit refspec (as `hermes update` now does), and when there is no
local branch yet create it at the fetched tip — checkout's own branch guess
ignores remote refs the configured refspec does not map. The remote's fetch
config is left as the user has it. Same change in install.sh and install.ps1.
2026-09-28 01:53:35 +05:30
Hermes Agent
c5380053b5 fix(install): keep the ffmpeg lock label, drop the network liveness test, document -SkipSetup
Review fix-ups on top of the re-pin:

- pm/lock.json: keep "version": "9.0.1". pm keys the store entry on
  `ffmpeg-<version>-<target>` and reinstalls on the artifact sha alone
  (pm/install.py::_identity / _entry_current), so the sha change already
  re-fetches the four BtbN targets. Bumping the label would also rename the
  macOS (martin-riedl, still 9.0.1) and Termux entries and re-stage unchanged
  bytes on every existing install for no reason.
- tests/pm/test_ffmpeg_pin_liveness.py: removed. It HEADs live GitHub URLs
  from the unit lane, and BtbN prunes dated autobuild tags after ~14 days
  (autobuild-2026-09-10-15-31 is already gone), so the test turns red for
  every PR on ~Oct 11 by construction. Upstream rot is what
  archive-inputs.yml (sha256 mirror on merge) and #122433 (re-pin on fetch
  failure) are for.
- website/docs/user-guide/windows-native.md + install.ps1 header: say that
  -SkipSetup is accepted as a deprecated alias for -NonInteractive instead of
  claiming it is rejected.

(cherry picked from commit def90331c2; ffmpeg lock/liveness hunks dropped, superseded by #125468)
2026-09-28 00:58:16 +05:30
Yuan Li
54421a43c6 fix(install): accept -SkipSetup alias and re-pin ffmpeg to a live BtbN tag (#125350)
The staged-installer rework (92686159d1) dropped the -SkipSetup switch
from install.ps1, so wrappers written against the old spelling
(hermes-desktop -SkipSetup -NonInteractive ...) die at parameter
binding with NamedParameterNotFound. Accept it as a deprecated alias
that folds into -NonInteractive.

pm/lock.json pinned ffmpeg artifacts to the dated BtbN autobuild tag
2026-09-10-15-31; BtbN prunes old dated tags, so fresh installs 404 on
GitHub and the sha256 mirror has nothing to serve (403). Re-pin all four
BtbN targets to a live tag (autobuild-2026-09-27-13-04, n9.0.2-12) with
digests taken from the GitHub release API.

(cherry picked from commit 91371d0b9a; ffmpeg lock/liveness hunks dropped, superseded by #125468)
2026-09-28 00:58:16 +05:30
teknium1
7605349f25 ci(install-e2e): run a path-filtered four-leg subset on pull requests
install-e2e.yml only ran on the clock, so nothing in front of a merge
installed a release and updated it on a real OS. A pull_request trigger,
path-filtered to the install/update surface (derived from 60 days of
update/install/pm commits), runs the new `pr` route of
generate-e2e-matrix.mjs with only the newest release tag sampled:

  linux   installer-script -> hermes-update (newest release -> PR)
  linux   installer-script -> hermes-update (PR -> NEXT)
  windows installer-script -> hermes-update (PR -> NEXT)
  macos   installer-script -> hermes-update (newest release -> PR)

The bundle-manifest validation job is skipped on PRs (bundled legs need
dispatch-only manifests). The full matrix stays on schedule and release.
2026-09-27 04:15:09 -07:00
teknium1
a3f454a287 fix(build): keep scripts/build/inputs.py free of pm imports; accept musl targets in its grammar
The Nix agent derivation builds scripts/build/*.py from a fileset that
does not include pm/, so importing pm.store.ALL_TARGETS there failed
nix flake check with ModuleNotFoundError. Extend the local target
regex with linux-(x64|arm64)-musl instead.
2026-09-27 03:24:39 -07:00
teknium1
842f162f7d fix(installer): refuse musl hosts without libstdc++ up front, naming it
PM's musl Node is the unofficial-builds musl archive, which links the
system libstdc++. On a stock Alpine (bash, git, curl) node and npm then
fail staged verification with raw relocation errors after the clone and
downloads. Check in the prerequisites stage and name the package.
2026-09-27 03:24:39 -07:00
teknium1
34d52883cb fix(installer): ELF interpreter decides musl in uv_bootstrap_target, as in pm/store
install.sh consulted ldd before the ELF interpreter while pm/store.py's
_is_musl_libc reads the native userland's ELF interpreter first, so a
glibc ldd (secondary toolchain, gcompat) could make the bootstrap stage
a different libc than PM later resolves. Read /bin/sh (then /bin/ls)
PT_INTERP first in both; ldd and the loader glob are fallbacks only.
2026-09-27 03:24:39 -07:00
teknium1
79e6f1960d fix(installer): restore executable bits on install.sh and gen-bootstrap-pins.py
The salvaged commits dropped both scripts from 100755 to 100644.
2026-09-27 03:24:39 -07:00
JoaoMarcos44
666c65c552 fix(build): accept PM musl targets in bundle inputs 2026-09-27 03:24:39 -07:00
JoaoMarcos44
87d1243623 fix(installer): detect musl uv when ldd is unavailable
Fall back to the musl loader if ldd cannot identify libc, while honoring
an explicit GNU libc report on hosts with a secondary musl toolchain.
Cover both paths against the pinned uv URL and digest.

Refs: #123682
2026-09-27 03:24:39 -07:00
JoaoMarcos44
24487f3db6 fix(pm): keep musl installs on compatible runtime artifacts
Skip glibc-linked FFmpeg on musl in both default install and update roots.
Select musl uv during the standalone shell bootstrap, and let the native
userland resolve libc before bootstrap Python build metadata.

Add regression checks for the closure, target precedence, and installer pins.
2026-09-27 03:24:39 -07:00
ymat19
5cd9dbc261 fix(install): recognize bare PATH= assignments and keep the appended line idempotent
wire_shell_path's existing-setup regex required a character before PATH=, so it missed bare assignments such as Fedora's ~/.bashrc (    PATH="$HOME/.local/bin:$HOME/bin:$PATH") and Debian's ~/.profile. The installer then appended its own line to .bashrc, .profile and .bash_profile, and because Fedora's .bash_profile sources .bashrc, login shells got ~/.local/bin on PATH several times.

Match bare assignments too, and make the appended line a no-op when PATH already contains ~/.local/bin.
2026-09-27 02:50:10 -07:00
teknium1
16da7f1b38 test(e2e/windows): real C:\Users profiles, serialized gateway phases, Git-for-Windows machines 2026-09-27 00:41:41 -07:00
teknium1
22fe26db2a test(e2e/windows): forward the windows_update opt-in knobs through run_tests.sh 2026-09-27 00:41:41 -07:00
liuhao1024
646c3c8ad5 fix(update): resolve npm's manifest through symlinks and fall back to the probe
npm_execpath can point through a symlink, and the manifest sits beside
the resolved CLI, never beside the link: resolve the realpath before
looking for package.json. Layouts without a readable manifest now fall
back to the pre-fix child probe instead of aborting with ENOENT.

Move the regression test to tests-js/node-deps.test.mjs (the module's
own suite) and cover the symlinked-execpath and fallback lanes there.
2026-09-26 23:48:52 -04:00
liuhao1024
9880fbb109 fix(update): read npm's version from its manifest instead of a child probe
The npm version probe spawned node-under-node before the reuse
short-circuit, so on Windows a Job-Object EBUSY spawn failure aborted
the whole dependency preparation even when the install was already
complete — leaving the pending-completion marker behind and turning
every launch into the same doomed completion pass (#123933).

npm's own package manifest states its version without any process
creation, so the probe lane can no longer fail: the install-receipt
key keeps its exact value, engine checks still run, and completed
installs reuse with zero spawns.
2026-09-26 23:48:52 -04:00
Brooklyn Nicholson
d06a3b8a54 fix(desktop-update): keep console selection from stalling the Windows hand-off
conhost blocks every write to a console while a selection is active. The
hand-off replays buffered child output through Write-HandoffLog after
`hermes update` exits, so a selection in a visible hand-off console held the
result, marker cleanup and relaunch until the user pressed Esc (#103222).

Turn QuickEdit off on the console input buffer for the run (restored on
exit), and skip the console echo while a selection is in progress. The log
file still gets every line.
2026-09-26 21:43:20 -05:00
Brooklyn Nicholson
21edd6d5ae fix(desktop): stop prescribing repair and antivirus review for an unverified update
A failed receipt check after a zero-exit update means the updated Desktop build
could not be read, not that the install is damaged. The old copy told users to
repair the installation and review antivirus quarantine, which is destructive
advice for a healthy install (#107685). Say what happened and give the one
non-destructive recovery step.
2026-09-26 21:43:07 -05:00
Hermes Agent
980318e689 fix(install): detect arch-suffixed desktop builds on repair/upgrade reruns (#94703)
electron-builder names the unpacked output <os>-unpacked on x64 and
<os>-<arch>-unpacked elsewhere (linux-arm64-unpacked, win-arm64-unpacked).
install.sh desktop_product_present and install.ps1
Test-DesktopProductPresent only listed the x64 names, so a rerun on an
ARM64 desktop install skipped the desktop rebuild and left a bundle built
from the previous code. List every unpacked dir main_desktop.py already
resolves, in both installers.
2026-09-26 20:29:47 -05:00
Brooklyn Nicholson
5a4ff55f43 fix(desktop-update): accept arch-suffixed unpacked dirs in the linux relaunch gate (#94703)
electron-builder names the unpacked dir linux-unpacked on x86_64 but
linux-<arch>-unpacked on every other arch (linux-arm64-unpacked is what
ARM ships). The gate hardcoded the x86_64 name, so a healthy ARM install
false-gated as "skew" on EVERY update, telling the user to reinstall an
app that was already correct. The ostree/symlink half was fixed earlier
(d3b090a3); this lands the remaining arch-dir half.

Resolve the unpacked dir the running binary actually lives in by scanning
the release dir's linux*-unpacked candidates (canonicalised both sides
before the compare, so the symlink fix's semantics are preserved, with a
first-found fallback so foreign targets keep gating as skew.

Tests drive the real --self-test-gate entry point; on BSD-readlink hosts
a PATH shim provides GNU  so the gate logic runs everywhere
(the existing linux_only symlink matrix covers -dependent paths).

Co-authored-by: C-Est-Dept <cestdept@example.com>
Co-authored-by: Sahilvishnaliya <sahil@example.com>
EOF
)
2026-09-26 20:29:47 -05:00
Hermes Agent
b24b8149dd fix(icons): give the dev Dock its own mac-grid png
Linux uses apple-touch-icon.png as the window icon and Nix requires it to
match the full-bleed launcher icon, so keep it full-bleed and point the
dev-only app.dock.setIcon at assets/icon-mac.png instead.
2026-09-26 18:10:54 -05:00
Hermes Agent
c8094dc399 fix(icons): put the dev Dock icon on the mac grid and enlarge the mac art
Dev runs replace the Dock icon with public/apple-touch-icon.png, which the
generator rendered full-bleed, so it drew ~24% larger than its Dock
neighbors. Render it from the mac-grid master like the icns targets.

On the 824 grid the plate matches peers, but the girl inside a white tile
with a ring read small; scale her 1.12x about the plate center for every
mac target.
2026-09-26 18:10:54 -05:00