Commit Graph

1950 Commits

Author SHA1 Message Date
ethernet
2ce138afb2 icon creation and freshenss check 2026-09-02 15:56:11 -04:00
ethernet
c9e1592622 fix(ci): restore acs timestamp on the msixbundle dlib sign + retry
The previous fix split the bundle sign into two passes (dlib sign with no
/tr, then a separate `signtool timestamp` against digicert). That was
wrong for MSIX: signtool silently exits 3 on an untimestamped
.msixbundle (appx signatures require a timestamp), and the ATS dlib
cannot parse a third-party timestamp server's response at all.

Restore the original single sign call with
`/tr http://timestamp.acs.microsoft.com /td SHA256` — the only
timestamp server the dlib can speak (electron-builder's default, and
what the build legs' .msix sign uses) — and wrap it in a 3-attempt
retry so acs's intermittent flakiness never fails the bundle (signtool
replaces the signature on re-sign, so a retry is safe).

Verified: node --check, 215 signing/r2/appinstaller tests pass; skill
reference corrected (MSIX bundle sign requires the acs timestamp;
digicert -> "no content extracted", no /tr -> silent exit 3).
2026-09-02 01:07:01 -04:00
ethernet
d8fa6c2463 fix(ci): sign the msixbundle envelope in two passes — dlib sign, then public RFC3161 timestamp
The publish-win32-updater job died at the bundle-envelope sign with
signtool exit 3 and "@url:`http://timestamp.digicert.com`: no content
extracted". Passing /tr on the SAME signtool call as /dlib hands the
timestamp URL to the Azure Trusted Signing dlib (`@url:` form), which
cannot extract a token from a third-party RFC3161 server.

Split it like the repo's own batch-sign-binaries.mjs (the pattern the
build legs use, and which the batch-sign tests pin):
1. `signtool sign /fd SHA256 /dlib <dlib> /dmdf meta bundle` — NO /tr.
2. `signtool timestamp /tr http://timestamp.digicert.com /td SHA256
   bundle` — NO dlib; signtool does the RFC3161 exchange itself (probed:
   digicert returns Status: Granted for a plain timestamp query).
3. `signtool verify /pa`.

The timestamp pass also retries (3 attempts) so a flaky server never
forces a re-sign of the ~2.7GB bundle.

Verified: node --check, 215 signing/r2/appinstaller tests pass.
2026-09-01 23:14:59 -04:00
ethernet
6c154fcb8b fix(desktop): make the desktop typecheck chain fully green
The electron tsconfig, renderer, e2e and electron-builder checkJs passes
now all pass cleanly. Three pre-existing issues fixed (the branch is
ahead of upstream/main, where none of these files exist):

- package-process-reap.ts: guard `isUnderInstallRoot` loop members with
  `typeof root !== 'string'` so the TS union (string | readonly string[])
  narrows before normalize() — the old `if (!root)` did not narrow arrays.
- msix-shared.mjs: add JSDoc param/return types to the four git-tag
  helpers and the stamp/XML/Content-Type helpers so the electron-builder
  checkJs pass (which walks this module via its import graph) stops
  flagging implicit-any.
- electron-builder.config.cjs: bind storeMsix to a local + a
  mustStoreMsix() assertion helper — `store` is typed boolean, so checkJs
  cannot correlate it with the optional storeMsix; the helper is the
  single documented assertion point (checkJs forbids `!`).
2026-09-01 22:53:18 -04:00
ethernet
a9793b3ea6 refactor(release): rename the nightly release channel to canary
The fast-moving desktop prerelease channel is now "canary" everywhere:
the tag shape (vX.Y.Z-canary.<ts>), the electron-updater/R2 feed dirs
(canary.yml / releases/<os>/canary/), the update-channel consts and CLI
choices, the MSIX build-number derivation, the App Installer channel
paths, and the Windows Store flight var (MS_STORE_CANARY_FLIGHT_ID).

Also renames the scheduled workflow to canary-release.yml and the
release test file to test_release_canary.py, and flips the CLI flags
(--canary / --prune-canaries / prune-canaries subcommand).

Unrelated "nightly" mentions are untouched: Brave's own browser channel
(browser_connect), cron scheduling prose (README, i18n, cron/browser/
kanban docs, zh-Hans), upstream skill docs (comfyui/unsloth/torchtitan),
evals fixtures, Node's node-nightly prereleases, and cron job names in
gateway tests.

Note: MS_STORE_NIGHTLY_FLIGHT_ID was renamed to MS_STORE_CANARY_FLIGHT_ID
in the workflow — the matching repo/org variable on GitHub must be
renamed in repo settings for the Store flight ring to keep working.
2026-09-01 22:15:17 -04:00
ethernet
65d0f03e52 ci(desktop): re-enable the Windows Store submission with nightly flights
Flip publish-win32-store back on (was a dummy skip), restore the store
variant build in build-win32, and wire the MSStore CLI submission:

- stable tags → production submission: msstore submission delete (clear
  any pending, tolerant) then msstore publish <Store-*.msixbundle> -id.
- nightly tags → package flight ring: msstore flights submission delete
  then msstore publish -f <flightId> -id. delete-then-replace so the
  newest nightly always wins the single-slot submission queue (chosen
  over skip-if-pending: always ship the newest, at the cost of cert
  churn).
- Gate: runs for stable when MS_STORE_PRODUCT_ID is set; runs for
  nightly only when MS_STORE_NIGHTLY_FLIGHT_ID is ALSO set, so the
  flight ring stays off until the flight exists in Partner Center.
- The r2 staging loop archives the per-arch Store-*.msix again (and the
  assembled universal bundle is archived by the store job).

Credentials (release-signing environment): MS_STORE_TENANT_ID,
MS_STORE_SELLER_ID, MS_STORE_CLIENT_ID, MS_STORE_CLIENT_SECRET
(secrets); MS_STORE_PRODUCT_ID, MS_STORE_NIGHTLY_FLIGHT_ID (vars).

Verified: yaml parses + needs graph resolves, store variant build
restored, bash branch harness (nightly→flight / stable→production)
executes the right msstore invocations, 215 js tests pass.
2026-09-01 21:30:23 -04:00
ethernet
c35a6d65c9 merge: upstream/main into ethie/pm-clean
Bring in 597 upstream commits while preserving the branch's intentional
divergence (pm store, MSIX desktop, Termux removal).

Resolution notes:
- local_runtime/local-models cluster (30 both-added files): took upstream's
  evolved version — our side was a stale feat-merge snapshot with zero
  post-merge commits, and our non-conflicted importers were verified against
  upstream's exports.
- install scripts (install.ps1/install.sh/setup-hermes.sh): kept our staged
  pm-store bootstrappers (upstream still ships the old monolithic installer).
- Deleted-by-us files (node-bootstrap.sh, install_ps1 tests, termux.md):
  kept deleted — the pm store replaced that machinery.
- runtime_repair.py: kept ours (pm-based), restored upstream's managed_uv.py
  which surviving upstream files still import.
- Core files (run_agent, update_cmd, main.py, hermes_constants, estop,
  aux_client, browser_tool, desktop entry, config_defaults): per-file merges
  combining both sides' features (upstream fleet-wide estop, PID-identity
  daemon kill, editable-install guard; our utf-8-sig sweep, stable channel,
  PYTHONPATH desktop entry).
- Desktop/i18n: took upstream's flag-gated local-models components; merged
  both sides' i18n keys; merged run-electron-builder.mjs and notifications
  tests keeping both sides' tests.
- pyproject.toml: upstream's expanded exclude-newer list + our
  google-cloud-pubsub entry; uv.lock regenerated from the merged pyproject.
2026-09-01 20:22:04 -04:00
Ben Barclay
180291162f feat(telemetry): opt-in shared-metrics exporter (#95278)
feat(telemetry): opt-in shared-metrics exporter
2026-09-02 08:35:36 +10:00
ethernet
f437776030 ci(desktop): disable the Windows Store submission path for now
App installer feeds are the update path; the MSStore flow is parked.
publish-win32-store is now a dummy skip (ubuntu notice step, kept in the
dependency graph so nothing downstream changes), and build-win32 no
longer builds the store variant (--variant=store removed). The r2
staging loop's Store-* archive branch and the store-bundle script stay
in the tree, documented, for re-enable.

Verified: yaml parses, needs graph resolves, no --variant=store remains
in the build step.
2026-09-01 18:13:47 -04:00
emozilla
43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
ethernet
5f3f9f3e72 ci(desktop): split release pipeline per-OS; gate MSIX work on windows only
The msixbundle + finalize jobs were both gated on the entire 6-leg build
matrix, so macos + linux legs blocked the win32 App Installer feed and
everything else downstream.

Restructure desktop-bundled-release.yml into per-OS jobs:

- build-win32 (x64 + arm64) — the only active builder legs; builds the
  bundled + Store variants, audits arch, uploads *.msix, stages to R2.
- build-darwin / build-linux — dummy skips for now (no macOS updater arm;
  linux unshipped). Matrix kept so downstream jobs stay green; re-enable
  by restoring the build body + runners.
- publish-win32-updater (was msixbundle) — needs build-win32 only; stages
  the App Installer feed (.appinstaller + universal .msixbundle).
- publish-win32-store (NEW, parallel) — bundles the two Store-*.msix into
  one universal Store .msixbundle and submits it to the Windows Store via
  the MSStore CLI (microsoft/microsoft-store-apppublisher@v1.4). Stable
  tags only; gated on MS_STORE_PRODUCT_ID var so it stays skipped until
  the release-signing environment is configured. The store bundle is left
  unsigned on purpose — Partner Center re-signs on ingestion.
- publish-darwin-updater (was finalize) — dummy skip; no mac feed to
  publish until the darwin electron-updater arm returns.

Shared plumbing: resolveWinSdkTools moves into msix-shared.mjs (single
resolver for both bundle jobs, kills the dead candidates var); new
bundle-store-msixbundle.mjs bundles the store per-arch packages and
prints the bundle path on stdout for the workflow.

Verified: node --check all scripts, yaml parses + needs graph resolves,
107 r2-release tests pass.
2026-09-01 15:53:55 -04:00
ethernet
375ce8eee5 ci: block tracked paths that collide case-insensitively
Linux is case-sensitive; Windows and macOS are not. Two tracked paths
differing only by case (README.md vs readme.md, src/Foo.py vs SRC/foo.py)
land fine on Linux and silently break every clone on a case-insensitive
host — the filesystem holds one, so checkout fails or whichever wins
clobbers the other. Git won't stop the pair from landing; it only warns
at checkout time on a case-insensitive FS. This is the enforcement point.

Adds scripts/check-case-collisions.py (index scan keyed on casefolded
full paths) + an unconditional workflow_call job wired into ci.yaml and
the all-checks-pass gate — unconditional because a collision can ship in
any kind of PR (docs, JS, config), not just Python, so gating on a
language lane would be the same passive-rule trap the infographic check
closes. Tests in tests/scripts/test_case_collision_check.py build
collisions via git update-index --cacheinfo so they run on
case-insensitive filesystems too.
2026-09-01 15:37:57 -04:00
Teknium
95d4265602 fix(install): provision supported Python from TUR on Termux and update docs 2026-09-01 11:25:10 -07:00
gustbr
8581120011 fix(install): enforce Termux Python upper bound 2026-09-01 11:25:10 -07:00
fangliquanflq
c454fbc5bd fix(installer): harden managed Python resolution 2026-09-01 09:28:01 -07:00
fangliquanflq
a2895e9968 fix(installer): preserve managed Python fallback version 2026-09-01 09:28:01 -07:00
fangliquanflq
60c7c0c667 fix(installer): require Hermes-managed Python on Windows 2026-09-01 09:28:01 -07:00
ethernet
9ed9cc8350 fix(desktop): ship chromium in the payload — drop stripFetchCache
stripFetchCache did two things: dropped fetch-<sha> download-cache dirs and
dropped the chromium-* store entries. Both were wrong for the current flow:

- fetch-* is now pruned by pm gc during the bundle (the payload and CI
  cache already ship only live store entries), so the function was belt
  and suspenders.
- chromium-* was DROPPED while signNestedChromium (the autosign pass
  right after it) was trying to SIGN it — the strip ran first, so the
  signing found nothing. Chrome's Mach-O stayed unsigned, Apple's notary
  rejected the app, and the shipped mac bundle had no browser at all.

Rip stripFetchCache out entirely: the payload keeps chromium, and the
darwin autosign pass now has the chromium trees to sign (--deep over the
.app, per-file over loose Mach-O) so notary accepts them. The Windows
payload's chromium binaries are covered by the batch Authenticode pass.
2026-09-01 12:26:35 -04:00
Teknium
28044757aa fix(desktop-update): self-heal a TCC-anchor-bricked venv in the posix hand-off (#95759)
Installs converted by the reverted macOS TCC interpreter anchor
(#95425/#95541) are left with a real-file venv/bin/python copy, a
.tcc-anchor-source marker, and python3/python3.N aliases that die at
interpreter init ("No module named 'encodings'"). venv/bin/hermes execs
venv/bin/python3, so EVERY CLI entrypoint is dead — hermes doctor and
hermes update included — and the desktop hand-off loops on "Update failed
(exit 1)" forever. No Python-side heal can ever run on this class; the
hand-off shell is the last surface that still executes, so the heal lives
there.

posix.sh gains, before the update invocation:

* tcc_anchor_heal — probe-gated (only fires when venv/bin/python3 fails
  a scrubbed-env `import encodings` boot probe), marker-validated
  (absolute path, outside the venv), staged with per-attempt backups and
  full rollback if the repaired interpreter still fails its probe.
  Two repair shapes:
  - alias-brick (#95541 class): the anchored copy boots — re-materialize
    python3* as REAL FILES of the anchor (hardlink/copy; an alias symlink
    onto the copy is the crash shape). Marker kept: this is exactly the
    layout ensure_tcc_anchor marks "active", so no anchor ping-pong.
  - full brick: restore python → symlink to the marker-recorded store
    interpreter (if it boots) and aliases → symlinks; marker removed.
    The unblocked `hermes update` then re-installs a boot-gated healthy
    anchor — one-shot convergence, not a loop.
  Fail-closed on missing/unbootable source (vanished uv store class),
  missing marker, or relative/in-venv marker paths.
* tcc_pick_update_invoke — if aliases stay dead but venv/bin/python
  boots (the launchd-gateway shape), drive the update via
  `venv/bin/python -m hermes_cli.main` instead of the dead hermes shim.
* Honest terminal message: an unrecoverable dead interpreter is reported
  as a venv repair problem instead of a generic "Update failed (exit 1)".
* --self-test-tcc-heal runs the real heal + invoke selection against an
  --install-root and reports, for the test harness.

Tests (tests/test_desktop_update_tcc_heal.py) drive the REAL posix.sh
functions on Linux against synthetic venv trees: healthy no-op, alias
heal, symlink restore, fail-closed classes, rollback on failed
verification, invoke fallback, and an A/B of the reported loop (bricked
venv/bin/hermes fails with the exact field error before, boots after).
Sabotage-verified (re-introducing the alias-symlink bug fails 2 tests).

NOT mac-live-tested (no macOS runner); the heal is platform-independent
shell exercised through the self-test path without the uname gate.

Recovery design (validation/staging/rollback/probe pattern) after
@aeonsong's #96231; in-update heal intent from @liuhao1024's #95775
(its target function no longer exists on main and its heal point is
unreachable on dead-CLI installs); heal-point and ping-pong analysis by
@ahrazzle and @tokenfires on #95759.

Fixes #95759
2026-09-01 08:36:03 -07:00
ethernet
029099b249 fix msix bundle signging with better timestamp sig 2026-09-01 10:32:12 -04:00
4dlt
3a7f2234a6 fix(cli): use Chromium's namespace sandbox when userns is available on Linux
The desktop launcher demanded a root-owned 4755 chrome-sandbox on every
Linux host and shelled out to sudo to configure it. Launched from the
.desktop entry there is no TTY, so sudo fails silently and `hermes
desktop` exits without a window — and every update rebuilds the helper
user-owned, re-breaking the app (#88032, #51327). The update hand-off's
relaunch gate blocked on the same condition, so post-update auto-relaunch
never fired either (#58593).

On hosts where unprivileged user namespaces work, Chromium uses its
namespace sandbox and never consults the setuid helper. Probe the actual
capability with `unshare --user --map-root-user true` (fails closed) and
skip the sudo path when the probe succeeds; hosts with userns disabled or
AppArmor-restricted (Ubuntu 23.10+) keep the existing setuid-helper and
--no-sandbox fallback behavior unchanged. Sandboxing stays fully enabled
in both cases.

Fixes #88032
Fixes #51327
Fixes #58593

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 02:32:39 -07:00
joaomarcos
c26f75baab fix(security): keep profile exports out of source and image contexts
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
2026-09-01 01:00:23 -07:00
ethernet
a9be133aea merge: local-models (upstream/feat/local-models) onto the pm-clean stack 2026-08-31 18:00:49 -04:00
ethernet
47f4ab3a17 feat(desktop): bundle, publish, and update the desktop app as MSIX
Wire the desktop app onto the pm store for real distribution:
- MSIX bundle: electron-builder config, appx assets, manifest, copilot
  key + deep-link routing, App Installer + Windows Store variant
  (sign only the msix; inner binaries covered by the package block map)
- Rust CLI shim (apps/desktop/shim) — bundled builds run from the store
  python + shim, never the venv; payload symlinks relativized so the
  relocatable venv survives relocation
- Cloudflare R2 release pipeline: publish binaries + update feeds,
  nightly channels/tags, stamp-first version resolution
- Update system: gate, uninstall steward, boot bootstrap, release
  channels, update receipts
- install.ps1 reduced to a 361-line stage-protocol bootstrapper (heavy
  deps are pm's job); darwin updater + update-channel mirror ripped
- doctor: main's re-landed TCC anchor kept, termux branches removed

Rebuilt from ethie/pm onto the pm-store stack. 22 hot files hand-merged;
uv.lock + package-lock.json keep main's newer dep tree; test_engines
reads the pm/lock.json pin; lazy_deps.py deleted (all 222 importers
migrated to pm in the foundation commit).
2026-08-31 18:00:48 -04:00
ethernet
f9656f111a refactor: remove Android/Termux support entirely
Rip out the Termux/Android install lane: constraints-termux.txt,
psutil_android extraction, termux-browser tests, the apt install-method
lane in install.sh, termux api detection, and the TUI termux layout.
Zero added termux lines; 255 removed. System prompt and skill-utils
termux branches drop with the platform they served.

Rebuilt from ethie/pm (f8236a2d91, 7ad9f0c7a5, 8446c871b6, 82fb8e09d1)
onto the pm-store stack. 11 hot files hand-merged; test_code_execution
restored to branch-final (termux path assertions removed).
2026-08-31 18:00:48 -04:00
ethernet
3d12e86ef1 feat(pm): unified package manager — pm store foundation
Introduce the pm store: a unified, hash-verified package store that
replaces lazy_deps and the old installer's ad-hoc tool downloads.
Store tools are provisioned on PATH (ffmpeg, node/npm via pinned uv),
with a resumable 8-way downloader, verify() returning failure reasons,
and adopt() made EPERM-safe. chromium ships in the payload for every
target. The 3600-line install.sh is replaced by a staged bootstrapper
(heavy deps are pm's job after this); setup-hermes.sh, Dockerfile and
nix pin tables are rewired onto the store. Old install-script tests,
lazy_deps/managed_uv/build_info, and the ps1/bash installer test
batteries are removed with the machinery they tested.

Rebuilt from ethie/pm onto upstream/main (ac6c8028e0) after the
utf-8-sig sweep. 16 hot files (main also churned them) hand-merged:
platform adapters, main.py, electron/main.ts, tui_gateway/server.py,
cua_backend, installer-tests workflow, install.sh (full rewrite),
setup-hermes.sh, plugins doc.
2026-08-31 18:00:48 -04:00
Shannon Sands
f5bb1e144d fix(gateway): address startup-watchdog review findings (OOF-298, PR #89750)
Independent review of the initial startup-liveness watchdog surfaced two
P1s and three P2s. All are addressed here.

P1 — legitimate slow startups (large state.db schema migrations inside
SessionDB.__init__, which run synchronously before the loop starts) could
exceed the fixed 300s deadline and restart-loop. The watchdog now checks
process CPU time (time.process_time(), process-wide) when the deadline
expires: continuous CPU consumption means a live migration, so the deadline
is extended (with a warning log per extension). The OOF-298 deadlock class
parks every thread in futex waits and accrues ~zero CPU, so it still fires
on schedule. Documented limitation: a spinning busy-wait deadlock reads as
progress and won't fire — the observed incident class is parked threads.

P1 — import-time deadlocks were outside coverage. The implementation moved
to a stdlib-only top-level module (hermes_startup_watchdog), and
hermes_cli/main.py arms it via an argv fast-path ("gateway" + "run" in
argv) BEFORE the heavy module-level import graph. gateway/startup_watchdog
remains as a re-export shim so the intuitive import path keeps working for
the disarm site, tests, and REPL use. Import-lightness is a correctness
property, tested via AST inspection: at fire time the wedged main thread
may hold the import lock, so the fire path performs no imports on its own
thread — the lifecycle-ledger write runs on a bounded-join helper thread
and os._exit happens regardless.

P2 — disarm/fire race: the handle now has an explicit state machine
(armed → disarmed | firing) guarded by a lock; whichever transition takes
the lock first wins, so a disarm landing after deadline expiry but before
the fire transition is honored. Regression test forces the exact
interleaving by blocking inside the CPU probe.

P2 — uncovered entry points: cli.py --gateway and scripts/hermes-gateway
run_gateway() now arm the watchdog before importing the gateway graph.
hermes_cli/gateway.py run_gateway() keeps an idempotent backstop arm for
programmatic callers.

P2 — respawn-storm backoff interaction: the storm breaker's intentional
backoff sleep (up to minutes, ~zero CPU — indistinguishable from a parked
deadlock) now calls kick_startup_watchdog(extra_s=backoff) so the deadline
is pushed past the sleep instead of firing mid-backoff.

Also: the faulthandler stack dump is now additionally written to
logs/gateway-startup-watchdog.log (stderr may be absent on detached/
windowless runs); the disarm site in gateway/run.py moved inside the
loop-confirmed branch (if the loop is NOT live, the milestone was not
reached and the watchdog must stay armed); hermes_startup_watchdog added
to pyproject py-modules so sealed venvs ship it; SERVICE_RESTART_EXIT_CODE
is duplicated in the stdlib-only module with a parity test against
gateway.restart.

Tests: 38 in tests/gateway/test_startup_watchdog.py (contracts incl.
stdlib-only AST check and shim re-export identity, config resolution,
arm/disarm/kick, CPU-progress extension vs no-progress fire, probe-failure
fails toward firing, disarm-vs-fire race, dump record + file stacks,
lifecycle ledger, custom exit code).
2026-08-31 14:01:39 -07:00
Futahua
58c1876d37 test(sessions): falsify the per-session fence across real processes
The unit tests share one interpreter, so they cannot exercise the failure the
fence exists to prevent: two SEPARATE gateway processes, each holding its own
snapshot of a conversation, both writing to it. That is how the defect was found
and it is the only way to show it is closed.

This drives two real `python -m tui_gateway.entry` processes over stdio and
checks the whole sequence, including the parts that are easy to get wrong:

  session.create claims nothing        an idle composer must not hold a session
  the lease keys on the STORED id      a lease keyed on the runtime handle would
                                       fence nothing, since two processes
                                       resuming one conversation have different
                                       runtime ids by construction
  B may still RESUME                   reading is never fenced; only writing is
  B's submit -> SESSION_NOT_OWNED      typed, and the registry is unchanged
  A killed, B retries -> accepted      a dead owner is pruned, not permanent

No provider is needed. The fence is checked before the agent is built, so a
submit that later fails for want of a model still proves who owns the session --
which keeps the probe free of credentials and of inference cost.

Against the parent commit it stops at the second check with an empty registry,
which is the defect stated exactly: with no cap configured, nothing was recorded
and therefore nothing could be refused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 12:36:33 -07:00
Teknium
29112bef09 chore: release v0.21.0 (2026.8.31)
Some checks are pending
Install & Update E2E / Pick release tags (push) Waiting to run
Install & Update E2E / update (push) Blocked by required conditions
Install & Update E2E / installer (push) Blocked by required conditions
2026-08-31 12:29:27 -07:00
kshitijk4poor
e60bd6aed2 chore: AUTHOR_MAP entries for fangliquanflq and sycamoregroupltd
Maps the contributor emails for the PR #99265 and #97779 salvages so
check-attribution passes on the salvage PRs.
2026-09-01 00:25:38 +05:30
semao0
39540a03fc fix(install): never adopt a pre-release Node.js build
install_node() picks the newest tarball out of
nodejs.org/dist/latest-v${NODE_VERSION}.x/ and installs it without ever asking
whether the binary inside is usable. That index currently serves
node-v26.8.0-<os>-<arch>.tar.xz -- a final-looking filename -- whose binary
reports v26.8.0-alpha.0.0.0. Node publishes the headers tarball named by
process.release.headersUrl only for final releases, so node-gyp cannot compile
against that build and every native module fails to install.

Probe the extracted tree before it replaces anything on disk, and fall back to
an older release line when the probe rejects it, instead of leaving the install
with an unbuildable runtime. Mirror the guard in node-bootstrap.sh, and let
_managed_node_tree_outdated() treat a pre-release tree as outdated so an
already-broken install heals itself -- the existing heal only fires below the
target major, and a pre-release sits above it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GrEaXSjvFoBXKxAHTjnUbS
2026-08-31 10:53:18 -07:00
fangliquanflq
12116f7619 fix(scripts): align retry recovery documentation 2026-08-31 10:05:13 -07:00
fangliquanflq
fd596e2492 fix(scripts): preserve update retry fallback 2026-08-31 10:05:13 -07:00
fangliquanflq
f10a231efa fix(scripts): clarify Windows update retry marker semantics 2026-08-31 10:05:13 -07:00
fangliquanflq
fd24ac94e0 fix(update): resume deferred Windows desktop updates 2026-08-31 10:05:13 -07:00
xxxigm
cffc8bb2a4 fix(install): stop a CLI install from building the desktop's node-pty
The browser-tools step ran a bare `npm install` at the repo root, which
resolves the root package.json's `apps/*` workspace glob. That materializes
apps/desktop and with it node-pty, which ships no Linux prebuild and falls
back to `node-gyp rebuild` — so the installer needs make/gcc on a machine
that will never launch Electron or a PTY addon. Since #85297 made a failed
npm install fatal, a host without a C toolchain (a stock CentOS/RHEL box,
for instance) cannot complete a CLI-only install at all; it just reports
"npm install failed or timed out".

Name the workspaces the install actually needs instead. ui-tui and web are
selected when present, with --include-workspace-root so the root's shared
ESLint devDependencies are not pruned by the scoped install — the same
closure `hermes update` already installs. A checkout with neither workspace
falls back to a root-only install, since npm fails hard on a workspace it
cannot find. Desktop dependencies keep coming from install_desktop(), which
is only reachable via --include-desktop.

Against a pristine tree the unscoped install reifies 1362 packages including
node-pty 1.1.0; the scoped one reifies 582 with no native desktop addon.

A fork force-push can 404 the compare API used by detect-changes, which
fail-opens with ci_review=true and blocks the PR on a ci-reviewed label
the install change does not need. Recover the file list from the pull
request files endpoint before that fail-open.
2026-08-31 10:00:55 -07:00
liuhao1024
cb89872c28 fix(install): clean up a broken managed Node and guard the termux probe
AI-review follow-up on #87467:
- On probe failure, remove the extracted ~/.hermes/node tree and the
  node/npm/npx bin links so later installer steps and retry runs start
  clean instead of resolving node to a binary that cannot start.
- The termux pkg branch had the same silent-success class: an empty
  version probe logged success and set HAS_NODE=true. Degrade with the
  binary's own error instead.
2026-08-31 10:00:46 -07:00
liuhao1024
8f8351aff3 fix(install): report a managed Node that cannot start, and preinstall libatomic1
install_node's post-install probe was
installed_ver=$(node --version 2>/dev/null) under set -e: when the
downloaded Node exists but cannot start (Node 26 linux-x64 builds link
libatomic.so.1, missing on minimal Debian/Ubuntu), the assignment
aborted the whole installer at exit 127 with the loader's explanation
discarded — installs died mid-sentence with no output at all (#87460).

- Probe now captures stderr and degrades with log_error carrying the
  loader message plus the libatomic1 hint instead of aborting.
- Debian/Ubuntu installs preinstall libatomic1 (best-effort, mirroring
  the existing apt idiom) so the common case just works.
- Termux branch's same-shaped probe gets a || true guard.
Fixes #87460
2026-08-31 10:00:46 -07:00
liuhao1024
0aa6b44917 fix(install): defer the partial clone's checkout so the throttle fallback engages
Review feedback on this PR: without --no-checkout, the blob fetch runs
inside git clone's own checkout step, so when the repo-scoped 429 hits
that fetch the whole clone exits non-zero, the else branch removes the
directory, and the fallback degrades to one more failed clone under
exactly the condition it exists for.

- Clone with --no-checkout (commits+trees only — small, passes the
  throttle); the blobs are then fetched by a separate 'git reset --hard
  HEAD' the retry can actually wrap. Verified on a local file://
  filtering remote: the no-checkout clone materializes nothing and the
  reset alone produces the full working tree.
- Fail closed: both reset attempts failing now removes the checkout and
  reports 'Failed to clone repository' instead of the previous '|| true'
  + unconditional clone_ok=true handing the installer a half-materialized
  tree printed as a success.
- The reset runs under a subshell cd so a failed materialization never
  leaves the shell in a deleted cwd, and the direct-retry loop bound now
  derives from $max_attempts (seq) instead of a hardcoded 1 2 3 4 that
  could drift from the reported attempt count.
2026-08-31 10:00:38 -07:00
liuhao1024
11afd07f16 fix(install): retry the HTTPS clone and degrade past repo-scoped 429s
GitHub throttles packfile generation for this repository with
repo-scoped HTTP 429s that are not client IP rate limits: an
anonymous clone of a small repo succeeds and the API quota is
untouched, but the single big pack behind --depth 1 dies
mid-transfer with 'RPC failed; HTTP 429 / expected packfile'. The
fresh-install clone path had no retry and no fallback, so a clean
machine exited 1 at the download stage and left a half-populated
install directory (same throttle as the update path in #89287).

Retry the HTTPS clone with linear backoff, removing the partial
clone between attempts; when every direct attempt fails, degrade
to a blobless partial clone and materialize the working tree with
a hard reset — many small packs instead of one big one, which is
what gets past the throttle. SSH-first ordering, the existing
installation update branch, and the commit-pin flow are unchanged.
2026-08-31 10:00:38 -07:00
emozilla
87f2f6733e Merge remote-tracking branch 'origin/main' into merge/local-models-main
# Conflicts:
#	website/docs/user-guide/features/memory.md
2026-08-30 00:28:50 -04:00
emozilla
b361fadd49 feat(local-runtime): derive the recommended model from memory size and bandwidth
The static recommended flag in catalog.json picked the dense 27B on
every machine, including unified-memory boxes where it decodes at
~13 tok/s while the 35B-A3B MoE does ~60. Replace the flag with a
per-machine derivation: decode is memory-bound, so predicted speed is
bandwidth over bytes-read-per-token, and the pick is the highest-quality
entry that runs resident and clears a 20 tok/s pleasant floor — else the
fastest resident entry, else the least-painful spill.

Catalog entries carry two authored fields in place of the flag:
quality (AA-informed ordering, editorially owned — never fetched at
runtime) and decode_fraction (share of weight bytes a token actually
reads; 1.0 dense, the active-slice ratio for MoE). The bandwidth axis
is the existing uma flag for now; measured per-machine bandwidth can
replace the class constants without touching the rule.

All three consumers derive: the pane badge and hero card through the
catalog route, quickstart's default target through the same resolver,
each gated on engine eligibility. The decision table lives on as a
checked-in test pinning every memory-class x bandwidth cell — a catalog
change flips cells in that file and the diff in review IS the editorial
sign-off. scripts/aa_quality_sync.py proposes quality updates at
authoring time; the commit decides.

The quickstart fixture's select_variant stub now constructs a real
VariantChoice — the resolver reads zero_spill, which the SimpleNamespace
stub lacked.
2026-08-29 22:38:40 -04:00
ethernet
6e294f543d fix: read source files with encoding=utf-8-sig
Read text files with the encoding utf-8-sig so a BOM at the start of a
file does not cause a Unicode decode error (Windows editors add BOMs).

Reconstructed from ethie/pm commits 48a32b135b + 013219e814 onto the
current upstream/main base: only the utf-8 -> utf-8-sig transforms were
carried (370 exact line pairs across 205 files); pm-rename hunks that
rode in the original commit were left to the pm-store commit, and
utf8sig hunks entangled with content changes ride their owning commit.

Rebuilt on ethie/pm-clean off ac6c8028e0 (upstream/main).
2026-08-29 21:23:00 -04:00
Teknium
0b10acc5b8 fix(install): keep install.ps1 pure ASCII — seed the SOUL text with '--' dashes
The synced identity text carries em-dashes, but install.ps1 must stay
pure ASCII (Windows PowerShell 5.1 reads BOM-less .ps1 in the ANSI code
page; a non-ASCII byte in a string literal desyncs the parser — see
tests/test_install_ps1_ascii_only.py, issues #66994/#67000). Seed the
ASCII-dashed variant there instead, and register that variant in
_LEGACY_TEMPLATE_SOULS so Windows installs converge onto the canonical
em-dash text on first run.
2026-08-29 18:10:47 -07:00
nftpoetrist
0610291b5a fix(prompt): sync DEFAULT_SOUL_MD with the #95681 identity rewrite
DEFAULT_AGENT_IDENTITY was rewritten in agent/prompt_builder.py (behavior
spec, exploration-thrift line deliberately removed) but the actual seed
written to disk on first run, hermes_cli/default_soul.py's
DEFAULT_SOUL_MD, was never updated. ensure_hermes_home() writes
DEFAULT_SOUL_MD into SOUL.md on every fresh install before the agent's
first turn, so virtually all real users end up as "SOUL.md users" seeded
with the pre-rewrite text -- including the exact "targeted and efficient
exploration" line the rewrite explicitly banned -- while the new
DEFAULT_AGENT_IDENTITY fallback essentially never serves the "fresh
install" audience its own PR body named as the target.

- DEFAULT_SOUL_MD now matches DEFAULT_AGENT_IDENTITY exactly.
- The pre-rewrite text is added to _LEGACY_TEMPLATE_SOULS so installs
  already seeded with it self-heal via the existing upgrade-in-place
  mechanism (same guarantee as the comment-only scaffold entries: the
  string carries zero user intent, so it's safe to replace).
- Synced the other places install.sh's own comment says "MUST match
  DEFAULT_SOUL_MD": scripts/install.sh, scripts/install.ps1,
  docker/SOUL.md, and the docs/i18n pages that quote the fallback text
  verbatim.
2026-08-29 18:10:47 -07:00
Mohamed HAMMANE
6da0ae1cf5 fix(install): preserve project config for locked uv sync (#82446) 2026-08-28 12:36:15 -07:00
Teknium
15eb5caf7a fix(install): Node 24.0–24.10 no longer passes the gates only to die at npm EBADENGINE
The locked dependency tree now carries @babel/* 8.x, which requires
node ^22.18.0 || >=24.11.0. Our engines.node arm said ^24.0.0 and the
installer gates (node_satisfies_build / Test-NodeVersionOk) accepted any
Node 24 — so a system Node 24.0–24.10 cleared every gate we own and then
failed 'npm install' with EBADENGINE under engine-strict=true.

- Raise the 24 arm to ^24.11.0 in root + desktop package.json and the
  package-lock.json mirrors
- Tighten node_satisfies_build (install.sh) and Test-NodeVersionOk
  (install.ps1) to 24.11+; update user-facing wording
- Add invariant tests: every engines.node arm floor must satisfy every
  locked dependency's engines.node, and the installer gates must encode
  the same floors as the manifest — so the next babel-style floor bump
  turns into a CI red instead of a user install outage
- docs: correct stale 'Node.js v22' provisioning claim
2026-08-28 12:20:40 -07:00
Jefferson Nunn
c9fa2bba45 fix(install): tier-0 locked sync no longer trips over UV_NO_CONFIG
The installer exports UV_NO_CONFIG=1 at script start (sudo -u hygiene,
#21269). That export also hides the project's own [tool.uv] policy —
exclude-newer and its package exemptions — from uv. The resolver then
runs under a different policy than uv.lock was resolved under, and
--locked makes that mismatch fatal:

  error: The lockfile at `uv.lock` needs to be updated, but `--locked` was provided.

Every fresh install hit this and fell through to the non-hash-verified
PyPI fallback tiers, defeating the point of Tier 0. Strip the variable
for this one invocation only; it stays exported for every other uv call.
Runtime code already strips UV_NO_CONFIG before its own locked syncs for
the same reason (hermes_cli/managed_uv.py).
2026-08-28 12:20:40 -07:00
Teknium
ad4adbbfeb fix(install): name the supported Node lines in the system-npm gate warning
Follow-up for salvaged PR #84397: Test-SystemNodeReady still said
'too old (Hermes requires Node >=22.22.0)' — Node 25 is not too old,
it's an unsupported line.
2026-08-28 05:12:33 -07:00
fangliquanflq
4d08f51581 fix(install): reject prerelease Node toolchains 2026-08-28 05:12:33 -07:00