Follow-up for salvaged #88472 (+ #100055 / #99995 / #88998 intent):
- add_contributor.py compares filenames with str.casefold(), the same key
scripts/check-case-collisions.py uses repo-wide, so non-ASCII folds
(ß ~ ss) are caught the way macOS/Windows fold them.
- The KNOWN_CASE_CONFLICTS allowlist is gone: the historical
agent@Agents-Mac-mini.local pair was removed on main (fcdae2cf0b), so the
repo-wide test asserts zero collisions.
- New tests: same-login different-spelling is still refused (the filename
pair is the problem, not the login) and the exact spelling stays
idempotent; casefold vs lower coverage.
Co-authored-by: alfred-amanda <288490622+alfred-amanda@users.noreply.github.com>
contributors/emails/ uses the email as the FILENAME, so two mappings differing
only in case are the same file on Windows and on default macOS. The tree has
such a pair today:
contributors/emails/agent@Agents-Mac-mini.local -> skip-agent
contributors/emails/agent@agents-Mac-mini.local -> momomojo
git writes one and then reports the other as modified in a FRESH clone, forever.
The repo cannot be checked out clean on those platforms, which breaks any tool
that gates on a clean tree -- our own Windows Desktop rebuild refuses with
"fresh clone is NOT clean" and never gets to build.
add_contributor() now refuses a mapping that case-collides with an existing one,
for the same reason it already refuses a conflicting login: the tool exists so a
typo cannot silently reassign commits, and a collision does exactly that on half
the platforms it lands on.
Two tests: the guard, and a directory-wide check that no NEW collision appears.
The existing pair is pinned in KNOWN_CASE_CONFLICTS rather than resolved here --
the two files name DIFFERENT logins, so picking one reassigns a contributor
commit history, and that is a maintainer call. Please resolve it; the pin keeps
the breakage visible and stops it spreading meanwhile.
Verified: pytest tests/scripts/test_contributor_map.py -- 9 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
Linux is case-sensitive; Windows and macOS are not. Two tracked paths
differing only by case (README.md vs readme.md, src/Foo.py vs SRC/foo.py)
land fine on Linux and silently break every clone on a case-insensitive
host — the filesystem holds one, so checkout fails or whichever wins
clobbers the other. Git won't stop the pair from landing; it only warns
at checkout time on a case-insensitive FS. This is the enforcement point.
Adds scripts/check-case-collisions.py (index scan keyed on casefolded
full paths) + an unconditional workflow_call job wired into ci.yaml and
the all-checks-pass gate — unconditional because a collision can ship in
any kind of PR (docs, JS, config), not just Python, so gating on a
language lane would be the same passive-rule trap the infographic check
closes. Tests in tests/scripts/test_case_collision_check.py build
collisions via git update-index --cacheinfo so they run on
case-insensitive filesystems too.
Installs converted by the reverted macOS TCC interpreter anchor
(#95425/#95541) are left with a real-file venv/bin/python copy, a
.tcc-anchor-source marker, and python3/python3.N aliases that die at
interpreter init ("No module named 'encodings'"). venv/bin/hermes execs
venv/bin/python3, so EVERY CLI entrypoint is dead — hermes doctor and
hermes update included — and the desktop hand-off loops on "Update failed
(exit 1)" forever. No Python-side heal can ever run on this class; the
hand-off shell is the last surface that still executes, so the heal lives
there.
posix.sh gains, before the update invocation:
* tcc_anchor_heal — probe-gated (only fires when venv/bin/python3 fails
a scrubbed-env `import encodings` boot probe), marker-validated
(absolute path, outside the venv), staged with per-attempt backups and
full rollback if the repaired interpreter still fails its probe.
Two repair shapes:
- alias-brick (#95541 class): the anchored copy boots — re-materialize
python3* as REAL FILES of the anchor (hardlink/copy; an alias symlink
onto the copy is the crash shape). Marker kept: this is exactly the
layout ensure_tcc_anchor marks "active", so no anchor ping-pong.
- full brick: restore python → symlink to the marker-recorded store
interpreter (if it boots) and aliases → symlinks; marker removed.
The unblocked `hermes update` then re-installs a boot-gated healthy
anchor — one-shot convergence, not a loop.
Fail-closed on missing/unbootable source (vanished uv store class),
missing marker, or relative/in-venv marker paths.
* tcc_pick_update_invoke — if aliases stay dead but venv/bin/python
boots (the launchd-gateway shape), drive the update via
`venv/bin/python -m hermes_cli.main` instead of the dead hermes shim.
* Honest terminal message: an unrecoverable dead interpreter is reported
as a venv repair problem instead of a generic "Update failed (exit 1)".
* --self-test-tcc-heal runs the real heal + invoke selection against an
--install-root and reports, for the test harness.
Tests (tests/test_desktop_update_tcc_heal.py) drive the REAL posix.sh
functions on Linux against synthetic venv trees: healthy no-op, alias
heal, symlink restore, fail-closed classes, rollback on failed
verification, invoke fallback, and an A/B of the reported loop (bricked
venv/bin/hermes fails with the exact field error before, boots after).
Sabotage-verified (re-introducing the alias-symlink bug fails 2 tests).
NOT mac-live-tested (no macOS runner); the heal is platform-independent
shell exercised through the self-test path without the uname gate.
Recovery design (validation/staging/rollback/probe pattern) after
@aeonsong's #96231; in-update heal intent from @liuhao1024's #95775
(its target function no longer exists on main and its heal point is
unreachable on dead-CLI installs); heal-point and ping-pong analysis by
@ahrazzle and @tokenfires on #95759.
Fixes#95759
The desktop launcher demanded a root-owned 4755 chrome-sandbox on every
Linux host and shelled out to sudo to configure it. Launched from the
.desktop entry there is no TTY, so sudo fails silently and `hermes
desktop` exits without a window — and every update rebuilds the helper
user-owned, re-breaking the app (#88032, #51327). The update hand-off's
relaunch gate blocked on the same condition, so post-update auto-relaunch
never fired either (#58593).
On hosts where unprivileged user namespaces work, Chromium uses its
namespace sandbox and never consults the setuid helper. Probe the actual
capability with `unshare --user --map-root-user true` (fails closed) and
skip the sudo path when the probe succeeds; hosts with userns disabled or
AppArmor-restricted (Ubuntu 23.10+) keep the existing setuid-helper and
--no-sandbox fallback behavior unchanged. Sandboxing stays fully enabled
in both cases.
Fixes#88032Fixes#51327Fixes#58593
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
Independent review of the initial startup-liveness watchdog surfaced two
P1s and three P2s. All are addressed here.
P1 — legitimate slow startups (large state.db schema migrations inside
SessionDB.__init__, which run synchronously before the loop starts) could
exceed the fixed 300s deadline and restart-loop. The watchdog now checks
process CPU time (time.process_time(), process-wide) when the deadline
expires: continuous CPU consumption means a live migration, so the deadline
is extended (with a warning log per extension). The OOF-298 deadlock class
parks every thread in futex waits and accrues ~zero CPU, so it still fires
on schedule. Documented limitation: a spinning busy-wait deadlock reads as
progress and won't fire — the observed incident class is parked threads.
P1 — import-time deadlocks were outside coverage. The implementation moved
to a stdlib-only top-level module (hermes_startup_watchdog), and
hermes_cli/main.py arms it via an argv fast-path ("gateway" + "run" in
argv) BEFORE the heavy module-level import graph. gateway/startup_watchdog
remains as a re-export shim so the intuitive import path keeps working for
the disarm site, tests, and REPL use. Import-lightness is a correctness
property, tested via AST inspection: at fire time the wedged main thread
may hold the import lock, so the fire path performs no imports on its own
thread — the lifecycle-ledger write runs on a bounded-join helper thread
and os._exit happens regardless.
P2 — disarm/fire race: the handle now has an explicit state machine
(armed → disarmed | firing) guarded by a lock; whichever transition takes
the lock first wins, so a disarm landing after deadline expiry but before
the fire transition is honored. Regression test forces the exact
interleaving by blocking inside the CPU probe.
P2 — uncovered entry points: cli.py --gateway and scripts/hermes-gateway
run_gateway() now arm the watchdog before importing the gateway graph.
hermes_cli/gateway.py run_gateway() keeps an idempotent backstop arm for
programmatic callers.
P2 — respawn-storm backoff interaction: the storm breaker's intentional
backoff sleep (up to minutes, ~zero CPU — indistinguishable from a parked
deadlock) now calls kick_startup_watchdog(extra_s=backoff) so the deadline
is pushed past the sleep instead of firing mid-backoff.
Also: the faulthandler stack dump is now additionally written to
logs/gateway-startup-watchdog.log (stderr may be absent on detached/
windowless runs); the disarm site in gateway/run.py moved inside the
loop-confirmed branch (if the loop is NOT live, the milestone was not
reached and the watchdog must stay armed); hermes_startup_watchdog added
to pyproject py-modules so sealed venvs ship it; SERVICE_RESTART_EXIT_CODE
is duplicated in the stdlib-only module with a parity test against
gateway.restart.
Tests: 38 in tests/gateway/test_startup_watchdog.py (contracts incl.
stdlib-only AST check and shim re-export identity, config resolution,
arm/disarm/kick, CPU-progress extension vs no-progress fire, probe-failure
fails toward firing, disarm-vs-fire race, dump record + file stacks,
lifecycle ledger, custom exit code).
The unit tests share one interpreter, so they cannot exercise the failure the
fence exists to prevent: two SEPARATE gateway processes, each holding its own
snapshot of a conversation, both writing to it. That is how the defect was found
and it is the only way to show it is closed.
This drives two real `python -m tui_gateway.entry` processes over stdio and
checks the whole sequence, including the parts that are easy to get wrong:
session.create claims nothing an idle composer must not hold a session
the lease keys on the STORED id a lease keyed on the runtime handle would
fence nothing, since two processes
resuming one conversation have different
runtime ids by construction
B may still RESUME reading is never fenced; only writing is
B's submit -> SESSION_NOT_OWNED typed, and the registry is unchanged
A killed, B retries -> accepted a dead owner is pruned, not permanent
No provider is needed. The fence is checked before the agent is built, so a
submit that later fails for want of a model still proves who owns the session --
which keeps the probe free of credentials and of inference cost.
Against the parent commit it stops at the second check with an empty registry,
which is the defect stated exactly: with no cap configured, nothing was recorded
and therefore nothing could be refused.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
install_node() picks the newest tarball out of
nodejs.org/dist/latest-v${NODE_VERSION}.x/ and installs it without ever asking
whether the binary inside is usable. That index currently serves
node-v26.8.0-<os>-<arch>.tar.xz -- a final-looking filename -- whose binary
reports v26.8.0-alpha.0.0.0. Node publishes the headers tarball named by
process.release.headersUrl only for final releases, so node-gyp cannot compile
against that build and every native module fails to install.
Probe the extracted tree before it replaces anything on disk, and fall back to
an older release line when the probe rejects it, instead of leaving the install
with an unbuildable runtime. Mirror the guard in node-bootstrap.sh, and let
_managed_node_tree_outdated() treat a pre-release tree as outdated so an
already-broken install heals itself -- the existing heal only fires below the
target major, and a pre-release sits above it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GrEaXSjvFoBXKxAHTjnUbS
The browser-tools step ran a bare `npm install` at the repo root, which
resolves the root package.json's `apps/*` workspace glob. That materializes
apps/desktop and with it node-pty, which ships no Linux prebuild and falls
back to `node-gyp rebuild` — so the installer needs make/gcc on a machine
that will never launch Electron or a PTY addon. Since #85297 made a failed
npm install fatal, a host without a C toolchain (a stock CentOS/RHEL box,
for instance) cannot complete a CLI-only install at all; it just reports
"npm install failed or timed out".
Name the workspaces the install actually needs instead. ui-tui and web are
selected when present, with --include-workspace-root so the root's shared
ESLint devDependencies are not pruned by the scoped install — the same
closure `hermes update` already installs. A checkout with neither workspace
falls back to a root-only install, since npm fails hard on a workspace it
cannot find. Desktop dependencies keep coming from install_desktop(), which
is only reachable via --include-desktop.
Against a pristine tree the unscoped install reifies 1362 packages including
node-pty 1.1.0; the scoped one reifies 582 with no native desktop addon.
A fork force-push can 404 the compare API used by detect-changes, which
fail-opens with ci_review=true and blocks the PR on a ci-reviewed label
the install change does not need. Recover the file list from the pull
request files endpoint before that fail-open.
AI-review follow-up on #87467:
- On probe failure, remove the extracted ~/.hermes/node tree and the
node/npm/npx bin links so later installer steps and retry runs start
clean instead of resolving node to a binary that cannot start.
- The termux pkg branch had the same silent-success class: an empty
version probe logged success and set HAS_NODE=true. Degrade with the
binary's own error instead.
install_node's post-install probe was
installed_ver=$(node --version 2>/dev/null) under set -e: when the
downloaded Node exists but cannot start (Node 26 linux-x64 builds link
libatomic.so.1, missing on minimal Debian/Ubuntu), the assignment
aborted the whole installer at exit 127 with the loader's explanation
discarded — installs died mid-sentence with no output at all (#87460).
- Probe now captures stderr and degrades with log_error carrying the
loader message plus the libatomic1 hint instead of aborting.
- Debian/Ubuntu installs preinstall libatomic1 (best-effort, mirroring
the existing apt idiom) so the common case just works.
- Termux branch's same-shaped probe gets a || true guard.
Fixes#87460
Review feedback on this PR: without --no-checkout, the blob fetch runs
inside git clone's own checkout step, so when the repo-scoped 429 hits
that fetch the whole clone exits non-zero, the else branch removes the
directory, and the fallback degrades to one more failed clone under
exactly the condition it exists for.
- Clone with --no-checkout (commits+trees only — small, passes the
throttle); the blobs are then fetched by a separate 'git reset --hard
HEAD' the retry can actually wrap. Verified on a local file://
filtering remote: the no-checkout clone materializes nothing and the
reset alone produces the full working tree.
- Fail closed: both reset attempts failing now removes the checkout and
reports 'Failed to clone repository' instead of the previous '|| true'
+ unconditional clone_ok=true handing the installer a half-materialized
tree printed as a success.
- The reset runs under a subshell cd so a failed materialization never
leaves the shell in a deleted cwd, and the direct-retry loop bound now
derives from $max_attempts (seq) instead of a hardcoded 1 2 3 4 that
could drift from the reported attempt count.
GitHub throttles packfile generation for this repository with
repo-scoped HTTP 429s that are not client IP rate limits: an
anonymous clone of a small repo succeeds and the API quota is
untouched, but the single big pack behind --depth 1 dies
mid-transfer with 'RPC failed; HTTP 429 / expected packfile'. The
fresh-install clone path had no retry and no fallback, so a clean
machine exited 1 at the download stage and left a half-populated
install directory (same throttle as the update path in #89287).
Retry the HTTPS clone with linear backoff, removing the partial
clone between attempts; when every direct attempt fails, degrade
to a blobless partial clone and materialize the working tree with
a hard reset — many small packs instead of one big one, which is
what gets past the throttle. SSH-first ordering, the existing
installation update branch, and the commit-pin flow are unchanged.
The synced identity text carries em-dashes, but install.ps1 must stay
pure ASCII (Windows PowerShell 5.1 reads BOM-less .ps1 in the ANSI code
page; a non-ASCII byte in a string literal desyncs the parser — see
tests/test_install_ps1_ascii_only.py, issues #66994/#67000). Seed the
ASCII-dashed variant there instead, and register that variant in
_LEGACY_TEMPLATE_SOULS so Windows installs converge onto the canonical
em-dash text on first run.
DEFAULT_AGENT_IDENTITY was rewritten in agent/prompt_builder.py (behavior
spec, exploration-thrift line deliberately removed) but the actual seed
written to disk on first run, hermes_cli/default_soul.py's
DEFAULT_SOUL_MD, was never updated. ensure_hermes_home() writes
DEFAULT_SOUL_MD into SOUL.md on every fresh install before the agent's
first turn, so virtually all real users end up as "SOUL.md users" seeded
with the pre-rewrite text -- including the exact "targeted and efficient
exploration" line the rewrite explicitly banned -- while the new
DEFAULT_AGENT_IDENTITY fallback essentially never serves the "fresh
install" audience its own PR body named as the target.
- DEFAULT_SOUL_MD now matches DEFAULT_AGENT_IDENTITY exactly.
- The pre-rewrite text is added to _LEGACY_TEMPLATE_SOULS so installs
already seeded with it self-heal via the existing upgrade-in-place
mechanism (same guarantee as the comment-only scaffold entries: the
string carries zero user intent, so it's safe to replace).
- Synced the other places install.sh's own comment says "MUST match
DEFAULT_SOUL_MD": scripts/install.sh, scripts/install.ps1,
docker/SOUL.md, and the docs/i18n pages that quote the fallback text
verbatim.
The locked dependency tree now carries @babel/* 8.x, which requires
node ^22.18.0 || >=24.11.0. Our engines.node arm said ^24.0.0 and the
installer gates (node_satisfies_build / Test-NodeVersionOk) accepted any
Node 24 — so a system Node 24.0–24.10 cleared every gate we own and then
failed 'npm install' with EBADENGINE under engine-strict=true.
- Raise the 24 arm to ^24.11.0 in root + desktop package.json and the
package-lock.json mirrors
- Tighten node_satisfies_build (install.sh) and Test-NodeVersionOk
(install.ps1) to 24.11+; update user-facing wording
- Add invariant tests: every engines.node arm floor must satisfy every
locked dependency's engines.node, and the installer gates must encode
the same floors as the manifest — so the next babel-style floor bump
turns into a CI red instead of a user install outage
- docs: correct stale 'Node.js v22' provisioning claim
The installer exports UV_NO_CONFIG=1 at script start (sudo -u hygiene,
#21269). That export also hides the project's own [tool.uv] policy —
exclude-newer and its package exemptions — from uv. The resolver then
runs under a different policy than uv.lock was resolved under, and
--locked makes that mismatch fatal:
error: The lockfile at `uv.lock` needs to be updated, but `--locked` was provided.
Every fresh install hit this and fell through to the non-hash-verified
PyPI fallback tiers, defeating the point of Tier 0. Strip the variable
for this one invocation only; it stays exported for every other uv call.
Runtime code already strips UV_NO_CONFIG before its own locked syncs for
the same reason (hermes_cli/managed_uv.py).
Follow-up for salvaged PR #84397: Test-SystemNodeReady still said
'too old (Hermes requires Node >=22.22.0)' — Node 25 is not too old,
it's an unsupported line.
Product-owner decision, 2026-08-27: the analytical need is stable
cross-window identity (retention curves, longitudinal install
behaviour), which the rotating pseudonym destroyed by design. The
feature has not shipped - zero consented users, zero production
transmissions - so identity semantics can change without breaking any
promise made to a user; existing (dev-only) consent windows carry
forward unchanged.
Removed in full rather than weakened in place:
- shared_metrics_identity.py (salt generation/rotation, HMAC-SHA256
derivation, payload substitution) and its 19-test file.
- The sender's derivation step. _freeze_identity keeps its validation
role (unreadable/non-object/id-less payloads still reject rather than
block the queue) and now records the raw install_id in
sent_install_id; _body rewrites the payload's install_id from that
frozen column, keeping byte-identical resends anchored to one
recorded value.
Consent surface updated in the same change: the setup wizard now states
plainly that packages carry the stable profile-scoped install ID (a
random UUID, no personal information, reset by deleting the
shared-metrics directory). No consent was ever collected under the old
wording in any shipped build.
Docs A.2/A.3 rewritten as decision records rather than silently
edited: A.2 records what is transmitted now and states the
consequences plainly (indefinite cross-package correlation is the
designed behaviour); A.3 records why rotation existed and why its
removal was accepted. The main-body "must not reuse the persistent
local identifier by default" escape hatch is exercised, not deleted:
that paragraph required exactly this product decision, which has now
been made. A.6's deletion note updated: install_id is now itself the
lookup key, so a future delete-on-request needs only a service-side
API, not a mapping.
Tests: the two privacy assertions invert deliberately
(test_the_stable_install_id_is_transmitted_as_is and the e2e wire
variant); freezing/byte-identical-retry coverage unchanged. Staging
E2E script now asserts transmitted == install_id.
258 targeted tests pass; ruff + footguns clean; both staging E2E
harnesses green with the raw id observed on the wire (202s).
Adds contributor email→username mappings via the new
contributors/emails/ system (one file per email) for two
compression contributors whose salvage PRs need attribution CI.
387 commits from main; no conflicts (verified with merge-tree before
merging). Overlap limited to hermes_cli/config_defaults.py and
hermes_cli/setup.py, both auto-merged; all shared-metrics surfaces
untouched by main.
Desktop builds legitimately produce no output for minutes on Windows
(the updater terminal is static until a step completes), so 300s risks
cancelling healthy long steps. 600s keeps the watchdog meaningful for
true stalls while clearing slow builds; HERMES_UPDATE_STEP_IDLE_SECONDS
still overrides. Per Teknium's review.
The cherry-pick landed the file with CRLF endings and a fixture
here-string whose col-0 brace prematurely terminated the handoff test's
SelfTest-block strip, tripping the drive-python-not-the-shim guard on
fixture code. LF restored (matching main), child-script loop inlined.
The #95625 watchdog cancels a step after StepIdleTimeoutSeconds (300s)
with no stdout/stderr. But a real `hermes update` is stdout-silent for
40+ minutes by design: the Electron/vite build streams to
logs/update.log, not the child's pipes (hermes_cli/update_cmd.py's
update-log tee). An output-only ceiling would therefore kill every
healthy large update at 5 minutes and mark it exit 124.
The drain now fingerprints logs/update.log (size + mtime) and, when the
idle ceiling is otherwise reached, treats growth of that file as
progress: reset the clock instead of terminating the tree. The stat
runs only once the ceiling fires, so the hot drain path never touches
the filesystem. HERMES_UPDATE_STEP_IDLE_SECONDS remains the override;
HERMES_UPDATE_PROGRESS_LOG points the self-test at its own file.
TDD proof: -SelfTestPipeDrain gains a fourth arm, logstall -- a step
that is silent on its pipes but appends to the progress log every
second and must reach its natural exit 3, never 124. Linux CI pins the
same contract at source level (TestIdleWatchdogCountsUpdateLogGrowth);
sabotage-verified: making the log-growth consult inert fails
test_stall_branch_consults_log_growth_before_terminating.