Commit Graph

2084 Commits

Author SHA1 Message Date
teknium1
a10620a669 fix(gateway): a record alone never means "attach", and --replace/--force work
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.

- ATTACH now requires a LIVE `identify` answer. The claim-time record is
  published with NO served set (the runner settles multiplex a moment later),
  and `host_gateway()` reports `served_known=False` when nothing answers. An
  owner whose served set is unknown yields a TRANSIENT refusal, never an
  attach: previously `default`'s claim published "default,other" before its
  socket bound, `other`'s systemd unit read that as "I am served", exited 78,
  and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
  instead of forcing `multiplex=True`, so a standalone gateway stops claiming
  the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
  `--force` into `_host_attach_or_none`. Both previously exited in the guard
  before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
  config refusal every supervisor parks on; "someone else serves me right now"
  is a runtime observation that ends when that process does. No unit files
  change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
  finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
  would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
  and re-enters with `replace=True`, so it can no longer attach to the corpse
  it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
  `st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
  it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
  generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
  `gateway status`/doctor across N profiles pays one probe, not N.

Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
2026-09-21 05:02:29 -07:00
teknium1
2463550c97 feat(website): plugin pages render the README from the pinned commit by default
The README section was gated on `readme: true` in the catalog YAML and no entry set it, so
all 222 plugin pages shipped without one. READMEs now render for every GitHub/GitLab entry
(fetched at the reviewed sha, subdir first then repo root, common casings and docs/README.md
as fallbacks); `readme: false` opts an entry out. Build proof: 222/222 READMEs rendered.
2026-09-21 01:15:33 -07:00
teknium1
10c273813d feat(plugin-catalog): screenshots and readme entry fields for the plugin pages
Two optional, submitter-controlled fields on a catalog entry feed the
entry's own page at /docs/plugins/<name>:

- `screenshots:` — up to 6 https URLs on GitHub hosts (same host rule as
  `image`, so the site never fetches from third-party hosts and a raw URL
  pinned to the sha is as immutable as the code).
- `readme: true` — the docs build renders the README from the PINNED
  commit (raw.githubusercontent.com / gitlab.com raw at <sha>), never live
  content, so what a user reads is what the reviewer read.

Validator rejects malformed values (admission), the loader parses and
drops off-host screenshots with a warning (client), and the extractor
emits `screenshots`, `readme`, `readmeUrl` and a `maintainerSlug` for the
author pages. Tests on all three.
2026-09-20 20:43:13 -07:00
liuhao1024
04845f5f3e fix(desktop): skip the local gateway restart on update when the Desktop is remote-served
A Desktop whose active connection is remote (SSH/remote/cloud, including the
registry primary) owns no local messaging gateway, yet the update hand-off
always ran `hermes update --gateway`. On hosts where launchd/service recovery
fails, the updater falls back to a detached local `gateway run --replace`;
with the same Telegram bot token as the remote VPS gateway, the two processes
compete for getUpdates and Telegram rejects one consumer, taking the
production bot offline (#117529).

Pass the ownership down the hand-off: globalRemoteActive() now adds
--no-gateway (posix) / -NoGateway (windows) when the Desktop is remote-served,
and both orchestrators drop --gateway from every update invocation (initial +
retry). The local-ownership default keeps --gateway exactly as before.
2026-09-20 19:30:16 -07:00
fangliquan
8c5a5deb79 fix(install): probe the command link directory for dependencies 2026-09-20 15:22:22 -07:00
joaomarcos
f58605b5c6 fix(desktop): keep Windows update hand-off hidden 2026-09-20 11:14:36 -07:00
liuhao1024
8828e356f7 fix(tests): let run_tests_parallel take the explicit file list from a file
--files carries the whole list as one argv element, and Linux caps a
single argument at MAX_ARG_STRLEN (128 KiB) - a much smaller limit than
ARG_MAX. The whole-suite list (~210 KB) dies with E2BIG in execve before
the runner's first line runs, so 'run the whole suite except one file'
cannot be expressed through --files at all.

Add --files-from PATH (or '-' for stdin), one path per line, mutually
exclusive with --files. A bare '-' after --files-from is normalized to
the '='-joined form because argparse treats '-' as a positional.
2026-09-20 10:51:40 -07:00
teknium1
6159bf4d87 runner: drop the RLIMIT_DATA worker cap
On the CI runner the pillow-heif HEIF encode in tests/tools/test_image_source.py hangs under
the cap (2/2 runs, 36/40 then 300 s timeout; 40/40 in 26 s on main). The wheel's encoder spins
on a failed allocation instead of erroring, so a heap cap on C code trades an OOM for a hang.
Keep the leak fix and the no-relaunch-on-kill rule.
2026-09-20 09:00:16 -07:00
teknium1
439ebe0ae9 fix(tests): stop the code_kernel reader-thread leak that OOM-killed test workers; cap worker heap
tests/tools/test_local_env_blocklist.py::TestPythonpathSelectiveStrip::
test_execute_code_composition_strips_inherited_hermes_entries hands code_kernel a MagicMock
process whose stdout/stderr only fake read(); the kernel drains with read1(), which on a bare
MagicMock never returns EOF. _stdout_reader died on `buf += chunk`, but _stderr_reader's
`while chunk := stderr.read1(4096)` spun forever appending mocks (each call growing
mock_calls) in a daemon thread that outlived the test: ~1 GB/min until the kernel killed the
worker. Five OOM incidents on this file (08-30, 09-13, 09-14, 09-16, 09-19), always blamed on
whichever test ran next. The fake now returns EOF from read1() as well.

Runner guardrails so the next runaway is a traceback, not a swap storm:
- each pytest worker runs under RLIMIT_DATA (8 GiB, Linux; HERMES_TEST_WORKER_MEM_GB, 0 = off).
  RLIMIT_AS is avoided on purpose: browsers spawned by tests reserve huge address space.
- a worker killed by signal or the file timeout is never --file-retries relaunched; a runaway
  relaunched while the first tree is still being reaped doubled the damage on 09-16.

Live: the file went from 20 min / 20 GB to 5.6 s / 140 MB. An allocate-forever probe dies with
MemoryError in 5 s; a SIGKILL'd worker launches once on this runner, twice on base.
2026-09-20 09:00:16 -07:00
teknium1
40bfcbda34 chore(lint): P33 — module-level constant frozen from a bridged HERMES_* env var
The advisory profile-scope lint had no pattern for the shape behind #115635 (a module
CONSTANT = _float_env(...)/os.environ.get("HERMES_...") of a var gateway/run.py bridges from
config.yaml). Fires on origin/main's `_RECONNECT_ATTENTION_AFTER_SECONDS`; 6 advisory hits on
head, all env-only knobs or already keyed per home (agent/redact.py).
2026-09-19 22:34:52 -07:00
teknium1
ece388eca8 fix: test-runner scratch root is /var/tmp/hermes-pytest; two tests follow TMPDIR
~/.cache/hermes-pytest failed seven files: a dot-dir ancestor made the hidden-dir search tests
see every fixture as hidden, and the longer root pushed the PulseAudio/voice AF_UNIX test
sockets past sun_path. /var/tmp is the FHS disk-backed temp root (never tmpfs), non-hidden,
and shorter than the old /tmp root. test_tool_result_storage asserts STORAGE_DIR instead of a
literal, and the zh-Hans bot-mode mirror follows the English code block it must copy.
2026-09-19 10:44:26 -07:00
teknium1
0746903679 fix: test-runner scratch lives in ~/.cache/hermes-pytest; env-less home fallback cannot raise
A runner root under HERMES_HOME/cache made conftest relocate every basetemp out of the live
Hermes home into ~/hermes-pytest-basetemp-*: deeper paths pushed AF_UNIX test sockets past
sun_path, and a profile home under $HOME renders as ~/… (which one test compared verbatim).
The root now sits in $XDG_CACHE_HOME/hermes-pytest (disk-backed, outside the Hermes home,
as short as the old /tmp root) and the test asserts the displayed form.

get_real_home()'s tempfile fallback raised RuntimeError on Windows when a child env carried
no HOME/USERPROFILE; it falls back to the old literal instead of crashing env construction.
2026-09-19 10:44:26 -07:00
teknium1
07df62d604 fix: sockets keep a short temp root; lint skips git-ignored artifacts; test runner scratch leaves /tmp
Chrome puts its SingletonSocket under $TMPDIR and AF_UNIX paths cap at 104/108 bytes, so a
deep HERMES_HOME (profile homes, test homes) made the new scratch TMPDIR kill Chrome at
startup ("Socket path too long") — two browser test files went red on the branch and green
on main. hermes_constants.socket_safe_tmpdir() keeps the scratch root when it fits and falls
back to the OS root for sockets only; the browser env and the code kernel RPC socket use it.

check_no_tmp_literals walked git-ignored runner artifacts (test_durations.json) and failed on
whatever the last test run wrote; it now skips `git ls-files --others --ignored` paths.

run_tests_parallel created its per-file temp roots in the system temp dir and exported no
TMPDIR, so a full-suite run wrote gigabytes of fixtures to tmpfs (3,225 leftover roots, 9.9 GB,
were sitting in /tmp on the dev box). Roots now live under HERMES_HOME/cache/scratch/pytest and
the test process inherits TMPDIR=<root>, so the existing cleanup removes every temp file.
2026-09-19 10:44:26 -07:00
teknium1
4586cde64d chore: mark the deliberate /tmp literals and shrink the lint baseline to one code block
The seventeen remaining literals are container-side paths, AF_UNIX socket-path-limit
candidates on darwin, detection needles, guard regexes and guidance text that tells the
model to avoid /tmp. Each carries an inline `no-tmp: ok — <why>` so the reason lives
next to the line; the baseline keeps only a fenced tree listing where a marker would render.
2026-09-19 10:44:26 -07:00
teknium1
e16fee4db1 refactor: resolve runtime temp paths via tempfile/TMPDIR instead of literal /tmp
Hermes now routes scratch space through HERMES_HOME/cache/scratch (exported as
TMPDIR), so every production path that still spelled out /tmp bypassed that and
kept teaching the agent the habit. Fallbacks in tool_result_storage,
code_execution_tool, process_registry, the ACP child HOME, mini_swe_runner's
local cwd, and the CI/profiling scripts now use tempfile.gettempdir(); shell
installers fall back to $TMPDIR (then HERMES_HOME) when mktemp is missing, and
repro/eval shells use `mktemp -d -t`. User-facing help text and sample payloads
(hermes send, approvals test, hooks test, voice-mode WSL hints, meet_bot debug
line) no longer suggest /tmp.

Container-side paths (mini_swe_runner docker cwd, sandbox base env, remote
sync tarballs) keep the literal because they name the sandbox filesystem,
not the host.
2026-09-19 10:44:26 -07:00
teknium1
3999096d18 ci: forbid literal /tmp paths outside a burn-down baseline
scripts/check_no_tmp_literals.py flags /tmp path tokens in production code, skills,
docs and prompt strings (tests, CI workflows, Dockerfiles, lockfiles, i18n mirror,
code comments and docstrings exempt; ${TMPDIR:-/tmp} idiom exempt). Opt out one line
with 'no-tmp: ok — <why>' on the line or the line above. _BASELINE lists pre-existing
hits per file: growth fails, burn-down is advisory (--strict-baseline / --print-baseline
to refresh). Wired into lint.yml next to check_compat_pointers.
2026-09-19 10:44:26 -07:00
teknium1
d0dbf2cbb6 fix(install): resolve uv shims before salvage and validate the copy in place
Install-Uv accepted any file at $HermesHome\bin\uv.exe, and copied whatever
`Get-Command uv` returned into that location. Chocolatey's bin\uv.exe is a
ShimGen launcher that locates ..\lib\uv\tools\uv.exe RELATIVE to itself, so
the copy is dead on arrival; `& exe --version` does not throw on a nonzero
exit, so the launcher passed the try/catch and the Python stage then failed
with "Python 3.11 not available" (#110350). The re-run path trusted the same
broken copy again.

Building on KoNit-K's Test-ManagedUvBinary and its three call sites:

- Test-ManagedUvBinary merges stderr, relaxes the error preference, and
  returns the `uv <version>` line only on exit 0 -- a launcher's error text
  can no longer surface as "Managed uv found (Cannot find file ...)".
- Resolve-UvShimTarget maps a candidate to the standalone binary before the
  copy: `<name>.shim` sidecar (Scoop), the Chocolatey bin\ -> lib\<pkg>\tools\
  layout, symlinks (winget Links\); other reparse points (WindowsApps
  app-execution aliases) have no copyable file and skip the salvage.
- The salvage rung validates the candidate where it lives, copies, then
  validates the COPY at its new location and removes it on failure, so the
  stage fails honestly instead of reporting success over a dead launcher.
- scripts/tests/test-install-ps1-uv-shim-validation.ps1 drives the real
  Install-Uv with compiled fake uv binaries (a working uv and a
  location-relative launcher) under stubbed installer rungs; wired into
  installer-tests.yml for pwsh 7 and Windows PowerShell 5.1.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-19 02:44:02 -07:00
KoNit-K
3662a1926d fix(install): validate every managed uv acceptance path
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-19 02:44:02 -07:00
KoNit-K
23298685f0 fix(install): reject broken copied uv shims 2026-09-19 02:44:02 -07:00
teknium1
7009611013 fix(desktop): document the afterExtract ordering, trim tests, refresh stale hook references
Follow-up to the salvaged #106846 commit (@JoaoMarcos44):

- after-extract.mjs: spell out WHY the stamp moved (electron-builder's
  beforeCopyExtraFiles rebuilds the PE with resedit for the ELECTRONASAR
  resource; rcedit then cannot commit to that exe, deterministically —
  #105629), and why disableAsarIntegrity was not taken.
- after-extract.test.mjs: two invariants — the hook wiring (afterExtract set,
  afterPack unset, ASAR integrity still on) and the stamp target
  (electron.exe on win32, nothing on other platforms). Red on origin/main.
- set-exe-identity.mjs / scripts/install.ps1: comments still named the
  afterPack hook / after-pack.mjs.
2026-09-19 02:31:27 -07:00
teknium1
c07708671d fix(gateway): every adapter session key goes through one seam (+ lint)
A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.

Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.

Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.

Phase 2 of #88715.
2026-09-18 22:04:43 -07:00
teknium1
0ddba07ad7 fix(ci): report an interpreter crash as CRASHED, not "no tests ran"
When a per-file pytest subprocess dies by signal (the sqlite cross-thread
close in #113186 was a SIGSEGV after every test had passed), faulthandler
prints "Fatal Python error: Segmentation fault" and no summary line, so
every count parses to 0. The runner filed that under "1 file where no
tests ran (collection/import error, ...)" beneath a summary that read
"0 failed" and exited 1 — two wrong diagnoses for one real bug, and it
was misread as a runner problem twice on main.

The runner now detects a signal death or a "Fatal Python error:" banner,
prefixes the captured output with the diagnosis (same convention as the
timeout path), marks the progress line CRASHED, counts "N files CRASHED"
on the summary line, lists the file in its own failure bucket, and no
longer trips the "NO TESTS RAN" guard for a crash that ran tests. The
flake retry already covers crashes (any non-zero rc), so nothing changes
there.
2026-09-18 19:41:51 -07:00
kshitijk4poor
8925c70a1c docs(whatsapp): group access section says what the gateway admits; env reference rows; trim bridge tests
Groups: policy, group-JID allowlist, and that participants are still authorised by
the gateway sender allowlist or pairing (`open` alone admits nobody without one);
`require_mention` defaults to false; WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
rows in the environment reference. The alt-id node tests collapse to one (the
participantAlt case duplicated the first-contact case; the "still resolves via mapping
files" case only re-asserted matchesAllowedUser).
2026-09-19 03:15:40 +05:30
Waldo
81fd9dc773 fix(whatsapp): honor group ingress policy in bridge
Use the configured group policy and group-JID allowlist at Node bridge intake instead of applying the DM sender allowlist to group participants.

Co-authored-by: Martin Gontovnikas <m@gon.to>
2026-09-19 03:15:40 +05:30
jinlingzi-cmd
3926c4209c fix: WhatsApp group messages dropped when LID sender has no lid-mapping (#72529) 2026-09-19 03:15:40 +05:30
Amit C
8a55373dbf fix(whatsapp): authorize first-contact LID senders 2026-09-19 03:15:40 +05:30
Andrey
b34ebc084a feat(send): add WhatsApp native mentions
Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
(cherry picked from commit ba7fd43826f9e36d886cc294e072378d5d082aa5)
2026-09-19 00:24:50 +05:30
teknium1
49397cf2b4 fix(tests): parallel runner reports a known flag's missing value as usage, not per file
The bare-flag check asked pytest's parser which tokens it does not know,
but wrapped parse_known_args in the same blanket except that guards
parser construction. A known flag with a missing value (`--tb` alone)
raises pytest.UsageError there, which the except turned into "nothing
unknown", so discovery ran and every per-file pytest died with
"argument --tb: expected one argument".

Keep the fallback around building the parser only; let parse_known_args
run outside it and surface UsageError (and the unknown-token list) as
this runner's own usage error before discovery. One invariant test.
2026-09-18 10:23:21 -07:00
teknium1
d7bedcee1e fix(tests): parallel runner rejects unknown bare flags with usage instead of sweeping
A bare token this runner does not own used to be forwarded to every per-file
pytest, so a typo (`--jbs`, or `--help` before #114065's fix) discovered the
whole suite and each file died with "unrecognized arguments" — hours to learn
about a typo. Validate the bare passthrough tokens against pytest's own
argparse parser (installed plugins loaded), and fail once with this runner's
usage (exit 2) before discovery. argparse handles the attached-short-value
(`-rA`), combined-flag (`-xvs`) and `-k expr` forms, so real pytest flags keep
passing through; tokens after a literal `--` are the caller's explicit choice
and are never validated. If pytest's parser cannot be built, the check is
skipped and behaviour is unchanged.

Follow-up to KoNit-K's `-h`/`--help` interception (#114065). Fixes #114059.
2026-09-18 10:23:21 -07:00
KoNit-K
41332e7851 fix(tests): handle parallel runner help flags 2026-09-18 10:23:21 -07:00
teknium1
3dcf0d49ad fix(install): let Rolldown name the missing binding; repair on every OS
The first cut derived the package from `binding-${platform}-${arch}` with an
exact-suffix match, which never matches Windows (`-msvc`) or Linux
(`-gnu`/`-musl`) names, so the repair only ever worked on macOS and
install.sh had to gate it there. Rolldown's own loader already resolves
platform, arch and libc and prints the exact `@rolldown/binding-*` it wanted
in its error chain; parse that instead and drop the gate. Also spawn npm
through a shell on Windows (Node refuses to spawn npm.cmd directly) and trim
the tests to the two invariants (no-op when it loads; installs exactly what
the loader asked for, then re-probes).
2026-09-17 00:25:57 -07:00
Gille
ae9f42accf fix(install): repair missing Rolldown bindings 2026-09-17 00:25:57 -07:00
Teknium
73521a8e37 fix(update): one bad workspaces glob no longer aborts the lockfile-churn cleanup
Path.glob raises NotImplementedError for a non-relative pattern, which a string `workspaces` (iterated char by char, so "/") or an absolute entry produces. The (OSError, ValueError, TypeError) catch missed it, so the error escaped to the caller's suppress(Exception) and no lock was reverted at all -- back to autostash every run. Non-list values are now ignored and each pattern is tried on its own so a bad one just owns nothing.

install.sh: read the workspace globs with `while read` instead of an unquoted $(...) so they are never pathname-expanded against the caller's CWD before `case` sees the pattern.
2026-09-16 17:44:36 -07:00
teknium1
ba153d6969 fix(install): installers keep the root lockfile when a workspace manifest is dirty
`scripts/install.sh::discard_update_lockfile_churn` and `scripts/install.ps1::Discard-LockfileChurn`
run the same per-directory predicate as `hermes update` did before the previous commit, so an
installer-driven update of a managed checkout (Desktop / bootstrap) reverted the root
`package-lock.json` whenever only `apps/desktop/package.json` was dirty, leaving spec and lock
out of sync for the next `npm ci`. Port the same ownership model: the root lock is kept when the
root manifest or any manifest matching a root `workspaces` glob is dirty; nested lockfiles are
still kept only with their sibling manifest; a manifest outside the graph still does not
protect the root lock.

install.sh reads the globs with sed/grep (no jq dependency) and matches with `case`; install.ps1
uses ConvertFrom-Json and `-like`. Bash side live-A/B'd in a throwaway repo (red on main, green
after; controls unchanged); the PowerShell side is the same shape and could not be executed on
this Linux host (no pwsh).
2026-09-16 17:44:36 -07:00
teknium1
659f5ae94f fix(update): keep .venv installs whole through ZIP fallback and venv repair
Follow-up to the cherry-picked #112966 so the uv-default `.venv` layout is
supported end to end, not only at the lookup sites:

- `_ZIP_PRESERVED_TOP_LEVEL` gains `.venv`. The dirty-tree guard runs
  `git status --ignored=matching`, so a gitignored `.venv/` surfaced as
  `!! .venv/` and refused every ZIP fallback on such installs ("the working
  tree has uncommitted changes or untracked files") — the live runtime was
  being treated as user data the overlay would destroy.
- `_repair_venv_on_current_checkout` recreates the venv at the resolved
  directory instead of a literal `venv`, so a broken `.venv` is rebuilt in
  place rather than growing a second environment that `project_venv_dir()`
  then prefers while `bin/hermes.cmd` still launches the old one.
- `_refuse_update_if_venv_foreign_owned` scans the resolved venv (the only
  remaining `PROJECT_ROOT / "venv"` literal on the update path).
- windows.ps1 names the actual shim path in the lock-timeout message.
- Tests: extend the real-git ZIP guard test with the `.venv` case (red
  before this commit); the holder-guard test now uses a kernel-runner child
  whose cmdline lacks `hermes_cli.main`, so only the venv-prefix arm can match
  it (red on origin/main); drop the `process.platform`-override vitest case,
  which exercised the same resolver as the `.venv` case with a different
  directory string.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-16 17:13:56 -07:00
KoNit-K
b6cc751e5f fix(update): support uv default venv in desktop updates 2026-09-16 17:13:56 -07:00
teknium1
720d06fd9e chore(whatsapp-bridge): refresh package-lock so the audit is clean
The committed lockfile still resolved body-parser 1.20.6 (nested qs 6.15.3),
express 4.22.2 with a top-level qs 6.15.3 and sharp 0.35.3, so the override
bump alone left `npm audit` at 3 findings (2 moderate qs, 1 high sharp) for
anyone installing from the lock — and `hermes doctor` kept flagging the
"WhatsApp bridge deps" row. `npm update --package-lock-only` inside the
existing manifest ranges: body-parser 1.20.8, express 4.22.3, qs 6.16.0,
sharp 0.35.4 -> `npm audit`: found 0 vulnerabilities. No manifest change
beyond the override bump; Baileys stays pinned at 7.0.0-rc13.

#109060 (Sep 12) was the earliest PR to move the override to 1.20.8 (it also
carried a redundant qs override, which 1.20.8 makes unnecessary).

Part of #112382

Co-authored-by: BenKalsky <1568840+BenKalsky@users.noreply.github.com>
2026-09-16 17:11:23 -07:00
Kevin Rajan
8ff91f6ac3 fix(whatsapp-bridge): bump body-parser override pin to 1.20.8
The 1.20.6 pin sat inside the vulnerable range it was meant to clear (1.20.5 - 1.20.6); 1.20.8 pulls qs ~6.16.0, clearing the transitive qs advisories.
2026-09-16 17:11:23 -07:00
fangliquan
1655dcd35d fix(compat): prune dependency trees from pointer scan 2026-09-16 16:58:02 -07:00
teknium1
034313e7cd feat: plugin catalog entries carry an optional version label and card image
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).

Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.

Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
2026-09-16 14:18:39 -07:00
teknium1
f0fd0650b5 test(update): autostash suite keeps the launchd restart scope off the host (#111866)
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
2026-09-15 21:49:33 -07:00
teknium1
f13a87e610 ci: advisory profile-scope pattern lint on the lines a PR adds
scripts/check_profile_scope_patterns.py runs the validated hazard regexes in
scripts/ci/profile_scope_patterns.json (18 of the 31 campaign patterns: every one has a scope_hint
and hits <= 50 sites on main; the wider ones are review greps, not lint) against the lines added
vs the PR base and prints file:line, pattern id/class and why. Always exits 0: most shapes have
legitimate sites (a standalone `hermes -p x` process where environ IS the profile), so the
reviewer reads each finding against its scope hint. Wired into lint.yml beside the public-surface
diff with continue-on-error.

Proof: the pre-fix tools/bot_relay.py (`env = dict(os.environ)`, before the served_profile_child_env
change) is flagged as P05/C2; the fixed file and this branch's diff vs main report 0 findings.
Test: a fixture with the hazard is flagged on the right lines, the scoped-builder version is not,
and the line filter hides hits outside the added range.
2026-09-15 10:59:22 -07:00
Robin Fernandes
d89cacc25f chore(free-tier): keep the rehearsal server out of the repo; the docs page explains the stand-in instead
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes
59fad62a40 fix(free-tier): review follow-ups — read the classifier's context, never replace a locked identity, re-inventory on retry
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
  the long-wait rate-limit check read the turn's extract_api_error_context()
  dict, which never carries welcome_refusal / welcome_route. They now read
  classified.error_context, where _nous_welcome_tier parks them; the guard
  records the classifier's reset_at. Tests drive the real classifier and the
  real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
  account (anon_account_locked) is now retired without replacement, matching
  the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
  re-inventories, so a provider connected during the cooldown keeps
  inference.
- The desktop's setup.ready listener only refreshes an untouched picker
  (oauth mode, no local endpoint, idle flow) and re-checks after the
  readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
  lock that _send re-acquires; the log is copied out first.

Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
  behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
  scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
  extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
  loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes
51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
teknium1
8a3cded09c fix(whatsapp): normalize device-qualified ids in every id comparison, bridge and Python
#89322 fixed the bridge-local normalizeWhatsAppId, but bridge.js has since
moved its id handling to bridge_helpers.js::normalizeWhatsAppId, which still
turned `<user>:<device>@lid` into the malformed `<user>@<device>@lid` for
mentionedJid / quoted participant / reaction keys, and the Python side
(gateway/platforms/whatsapp_common.py::_normalize_whatsapp_id) did the same
':'->'@' swap on botIds. Drop the local duplicate in bridge.js, import the
helper, and strip the `:<device>` suffix on both layers so the bot's own ids
compare equal to the bare ids WhatsApp sends for mentions and quotes.

One invariant test: device-qualified botIds match a bare mentionedId and a
bare quotedParticipant; a plain group message still does not trigger.
2026-09-15 04:41:35 -07:00
ebs
6bf1032609 fix(whatsapp): strip :<device> suffix in normalizeWhatsAppId
normalizeWhatsAppId did String(value).replace(':','@'), which turns a device-qualified id
like '116342762025117:14@lid' into the malformed '116342762025117@14@lid'. The bot's own id
(sock.user.id / sock.user.lid) carries the :<device> suffix while inbound mentionedJid and
contextInfo.participant (quoted message author) do not, so the bot's id never matches its
botIds set -> @mention and reply-to-bot are never detected in groups. Strip the :<device>
suffix instead so all id forms compare consistently.
2026-09-15 04:41:35 -07:00
teknium1
d228013832 fix(ci): feed the timeout scaler only healthy durations, and wire the cache in CI
Greptile's two findings on the original PR were both right.

1. The scaler read test_durations.json from the checkout, but CI ran on
   a fresh runner where that file never exists (it is gitignored and the
   slicing-era artifact/merge job that produced it is gone). The feature
   was inert exactly where the false FLAKY kills happen. tests.yml now
   restores the most recent main-saved cache before the run (PRs read
   only) and saves it after a green push to main, mirroring the
   ci-timings-baseline restore/save pattern already in ci.yaml.

2. _save_durations persisted every file's total subprocess wall,
   including the ~cap of a timed-out attempt and the retry-summed wall
   of a FLAKY file. With the scaler that compounds: a hang cached at
   ~300s earns 900s next run, then ~900s cached earns 2700s, until the
   job timeout is the only bound. _clean_pass_durations drops failed and
   FLAKY files from the write so a file's cached duration is always a
   first-attempt-clean measurement; those files keep their previous
   known-good entry.

Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
2026-09-15 03:47:55 -07:00
Teknium
9a69785790 fix(ci): scale per-file test timeout by cached duration to stop false FLAKY kills
The flat 300s --file-timeout SIGKILL'd known-slow large-collection
files when CI load dilated their runtime past the cap; the automatic
one-shot retry then passed, manufacturing a FLAKY report for a healthy
file. Seen 2026-08-18 on main run 32155223248's sibling PR runs:
tests/test_hermes_state.py (239 tests) killed at 300s on attempt 1,
passed in 205s on retry.

_effective_file_timeout() now gives each file
max(flat_cap, 3 x last cached duration) from test_durations.json.
The bound is only ever raised — genuinely hung files are still killed,
uncached files keep the flat cap, and --file-timeout/HERMES_TEST_FILE_TIMEOUT
semantics are unchanged.

Includes a sabotage-verified unit test (fails without the scaler).
2026-09-15 03:47:55 -07:00
teknium1
087c4f4f4e fix(tests): drop the RLIMIT_AS memory cap; the leak sweep is the fix
The per-process address-space cap was defense-in-depth on top of the
SessionDB leak sweep (nobody sets the knob; the sweep removes the leak).
On the 96-worker CI runner it was also the only PR-specific difference
when tests/tools/test_image_source.py hung to the 600s SIGKILL while the
same file passes in ~30s on every sibling branch: RLIMIT_AS counts virtual
reservations, and image/threading libraries reserve far more address
space than they touch. Keep the sweep, remove the cap and its env knob.
2026-09-15 03:44:50 -07:00