Commit Graph

2037 Commits

Author SHA1 Message Date
teknium1
d228013832 fix(ci): feed the timeout scaler only healthy durations, and wire the cache in CI
Greptile's two findings on the original PR were both right.

1. The scaler read test_durations.json from the checkout, but CI ran on
   a fresh runner where that file never exists (it is gitignored and the
   slicing-era artifact/merge job that produced it is gone). The feature
   was inert exactly where the false FLAKY kills happen. tests.yml now
   restores the most recent main-saved cache before the run (PRs read
   only) and saves it after a green push to main, mirroring the
   ci-timings-baseline restore/save pattern already in ci.yaml.

2. _save_durations persisted every file's total subprocess wall,
   including the ~cap of a timed-out attempt and the retry-summed wall
   of a FLAKY file. With the scaler that compounds: a hang cached at
   ~300s earns 900s next run, then ~900s cached earns 2700s, until the
   job timeout is the only bound. _clean_pass_durations drops failed and
   FLAKY files from the write so a file's cached duration is always a
   first-attempt-clean measurement; those files keep their previous
   known-good entry.

Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
2026-09-15 03:47:55 -07:00
Teknium
9a69785790 fix(ci): scale per-file test timeout by cached duration to stop false FLAKY kills
The flat 300s --file-timeout SIGKILL'd known-slow large-collection
files when CI load dilated their runtime past the cap; the automatic
one-shot retry then passed, manufacturing a FLAKY report for a healthy
file. Seen 2026-08-18 on main run 32155223248's sibling PR runs:
tests/test_hermes_state.py (239 tests) killed at 300s on attempt 1,
passed in 205s on retry.

_effective_file_timeout() now gives each file
max(flat_cap, 3 x last cached duration) from test_durations.json.
The bound is only ever raised — genuinely hung files are still killed,
uncached files keep the flat cap, and --file-timeout/HERMES_TEST_FILE_TIMEOUT
semantics are unchanged.

Includes a sabotage-verified unit test (fails without the scaler).
2026-09-15 03:47:55 -07:00
teknium1
087c4f4f4e fix(tests): drop the RLIMIT_AS memory cap; the leak sweep is the fix
The per-process address-space cap was defense-in-depth on top of the
SessionDB leak sweep (nobody sets the knob; the sweep removes the leak).
On the 96-worker CI runner it was also the only PR-specific difference
when tests/tools/test_image_source.py hung to the 600s SIGKILL while the
same file passes in ~30s on every sibling branch: RLIMIT_AS counts virtual
reservations, and image/threading libraries reserve far more address
space than they touch. Keep the sweep, remove the cap and its env knob.
2026-09-15 03:44:50 -07:00
Teknium
d070e480a3 fix(tests): close leaked SessionDB handles suite-wide and cap pytest memory
Root cause of the 2026-08-16 OOM incidents (three runs of
`python -m pytest -o addopts= -q tests/hermes_cli/` ballooning to
16-25 GB RSS and getting killed): ~40 files under tests/hermes_cli/
construct SessionDB() directly and never close it. Each instance keeps
the writer connection (state.db + -wal fds), up to _READ_POOL_MAX pooled
readers with their SQLite page caches, and — once token accounting has
run — an atexit registration that pins the instance alive until
interpreter exit. In one process over 637 files those accumulate without
bound; the sanctioned per-file runner masks it, so CI never saw it.

Fix the class, not the sites:

* hermes_state: register every successfully constructed SessionDB in a
  test-only WeakSet (populated only when HERMES_TEST_ISOLATION is set,
  i.e. under this test suite; production never touches it).
* tests/conftest.py: autouse _close_leaked_session_dbs teardown closes
  everything left in the registry after each test. close() is idempotent
  and unregisters the pinning atexit hook, so instances become
  collectable.
* tests/conftest.py: session-scoped _pytest_memory_cap applies a
  defensive RLIMIT_AS of 12 GiB (Linux only) so any future in-process
  leak fails fast with MemoryError instead of eating the box.
  Overridable/disable-able via HERMES_PYTEST_MEM_CAP (documented in
  scripts/run_tests_parallel.py).
* tests/hermes_state/test_session_db_leak_sweep.py: behavior contract
  for registration, idempotent close, and the cross-test sweep.

Measured (capped single-process `pytest -o addopts= -q tests/hermes_cli/`):
peak RSS 4.16 GiB before -> 1.67 GiB after; per-test open .db fd count
previously climbed monotonically (0 -> 12 -> 17 -> 104 within the
SessionDB-heavy files), now stays bounded (<= 5, transient). Sanctioned
runner over the affected 35 files: 495 passed, 0 failed, no FLAKY.

Incident evidence: ~/.hermes/logs/oom-incidents/20260816-202114
(fd dumps show 100+ open state.db/state.db-wal handles across pytest
tmpdirs; 3rd recurrence that day).
2026-09-15 03:44:50 -07:00
teknium1
2de17e5d40 feat(plugin-catalog): default shelf is Desktop, catch-all is General; categorise today's six entries
Teknium's call: most community submissions are Desktop panes, so an entry
without a category lands on the Desktop shelf; "other" becomes "general" for
plugins that genuinely span areas. Shelf order puts Desktop first. The six
entries merged today (pets-all, newswire, auto-titler, live-voice,
metamask-wallet, web-octen) get explicit categories.
2026-09-14 21:00:29 -07:00
teknium1
55dbd7f6e1 feat(plugin-catalog): shelve the catalog by category (Memory, Desktop, Platforms, …)
The catalog page was one undifferentiated grid filtered only by tier, so a
memory provider sat between two Desktop panes. Entries now carry an optional
``category`` (memory | desktop | platform | web | tools | voice | automation |
models | other, default other) that the loader, the admission validator and
the site extractor all understand.

/docs/plugins renders one shelf per category in browse mode, a category pill
row under the tier pills, a clickable category chip on every card, and a
results bar (active category, count, clear) when a filter or search flattens
the view. ``hermes plugins catalog`` gains a Category column and groups by it.
All 18 shipped entries are categorised. Unknown categories fail admission
(same contract as tier) so a typo cannot create a phantom shelf.
2026-09-14 21:00:29 -07:00
teknium1
49c6d4a9e0 test(contracts): tests mirror tui_gateway/; the runtime-artifact spoof test asserts the new 4000
tests/contracts -> tests/tui_gateway/contracts (tree-layout rule: tests mirror a source
package). test_rpc_params_cannot_spoof_runtime_artifacts: forged owner_transport /
owner_session_record / owner_token keys are now refused at the wire (4000 + key path)
instead of silently dropped before the handler; the invariant (no steer reaches the
agent) is unchanged and asserted directly.
2026-09-14 06:12:19 -07:00
teknium1
cdf949877a fix(contracts): generator emits prettier-style TS directly (no Node in the Python CI lane)
The staleness test regenerates in the Python lane, which has no
node_modules; prettier-dependent output would make the check pass locally
and fail in CI (or the reverse). Single-quoted literals, bare identifier
keys, no trailing commas or whitespace — prettier --check is clean on the
committed file.
2026-09-14 06:12:19 -07:00
teknium1
0250c8bcae feat(contracts): declare every gateway method, server request and event; commit the generated TS + OpenRPC (#110522, part 2)
215 methods, 13 server→client requests and 67 notifications now have Pydantic
contracts under tui_gateway/contracts/<topic>.py, rendered to
apps/shared/src/gateway-contract.generated.ts (616 types) and
gateway-contract.openrpc.json. tests/contracts/test_generated.py pins both
files to an in-memory regeneration and asserts catalog completeness from the
CODE side (every registered handler / emitted event / sent request has a
contract, nothing orphaned). scripts/ci/classify_changes.py runs the Python
lane when either generated file changes.

Phantom fields the hand-typed TS carried and no emitter ever set:
tool.start.todos, error.reason, voice.transcript.voice_stopped.
2026-09-14 06:12:19 -07:00
teknium1
d4840f9236 feat(contracts): Pydantic wire-contract registry, runtime validation and TS/OpenRPC generator (#110522, part 1)
tui_gateway/contracts/ is the single source of truth for the JSON-RPC wire:
Params/Result/Payload bases, a registry of METHODS / SERVER_REQUESTS /
EVENTS, runtime validation (unknown/mistyped params answer 4000 with the
field path; a result or payload that violates its model raises under the
test suite and logs once in production), and the server-request contracts
+ shared value shapes as the authoring template. scripts/gen_gateway_contracts.py
renders the tables through a small JSON-Schema-subset walker into
apps/shared/src/gateway-contract.generated.ts and
gateway-contract.openrpc.json (unsupported constructs raise at generation).
Method contracts for the 213 remaining handlers follow in the next commits.
2026-09-14 06:12:19 -07:00
teknium1
d10bb2ab6f test: make tests/ mirror the source tree; drop issue numbers from filenames
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.

Parallel directories for one source package, folded into the mirror:
  tests/acp        -> tests/acp_adapter   (its __init__/conftest move with it)
  tests/cli        -> tests/hermes_cli    (prompt_toolkit fixture merged into
                                           hermes_cli/conftest.py)
  tests/run_agent  -> tests/agent         (backoff fixture becomes
                                           agent/conftest.py)
  tests/relay      -> tests/gateway/relay
  tests/state      -> tests/hermes_state

246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.

Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.

Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).

Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
2026-09-13 09:18:02 -07:00
teknium1
e06efbcff0 fix(ci): desktop slash-registry dump gets --check; apps/ contract JSON edits run the Python lane
scripts/dump_desktop_slash_registry.py --check exits 1 (no write) when the
committed apps/desktop/src/lib/desktop-slash-registry.json differs from
desktop_surface_registry(); the docstring now cites the test that actually
fails (test_desktop_slash_registry.py, not test_commands.py). The pytest
runs --check as a subprocess (green on the committed file) and calls
check() on a hand-edited copy (red).

scripts/ci/classify_changes.py: _PY_SKIP includes apps/, so an apps-only
edit of that JSON (or apps/shared/src/gateway-events.json) skipped the
Python lane and the only equality check never ran in CI. A table of
cross-language contract files now forces python:true; two classifier rows
pin it.
2026-09-13 06:50:57 -07:00
teknium1
458595a20b refactor(desktop): slash block-list derives from the Python command registry; only 5 TS-only names stay hand-typed
34 of the 46 `NO_DESKTOP_SURFACE` rows in desktop-slash-commands.ts were a
byte-for-byte copy of `desktop=` on the matching CommandDef in
hermes_cli/commands.py (7 more were aliases of those rows). The live
`commands.catalog` already carries that metadata; the static list was the
offline fallback and would silently drift on the next registry edit.

Now `hermes_cli/commands.py::desktop_surface_registry()` is the one author
of `/name -> desktop` (aliases included). `scripts/dump_desktop_slash_registry.py`
writes it to apps/desktop/src/lib/desktop-slash-registry.json, which the
desktop imports as its offline fallback (`registryUnavailableSpecs`). Five
names the Python registry has never heard of stay in an explicit
`TS_ONLY_NO_DESKTOP_SURFACE` with the reason WHY: `/density /details /logs
/mouse` are Ink-process-local display toggles (handled in
ui-tui/src/app/slash/commands/core.ts; advertised via `_TUI_EXTRA`), and
`/pets` is the plural typo of the desktop's own `/pet` action. `/switch`
was never a block-list row (it is a `/resume` alias) — not a finding.

Cross-language contract: tests/hermes_cli/test_desktop_slash_registry.py
asserts the committed JSON == desktop_surface_registry() and that every
alias carries its canonical value; desktop-slash-commands.test.ts asserts
every dumped row is unavailable/unsuggested offline with the dumped reason
and that the TS-only set is disjoint from the dump. Both sides fail on
drift (sabotage: flipping one `desktop=` in commands.py -> Python test
"stale"; adding `/clear` to the TS-only list -> vitest disjointness fails;
dropping `registryUnavailableSpecs()` -> 4 existing vitest cases fail).

Behavior change: none for users. `/model` (`desktop="hidden"`) keeps its
local picker spec; `hidden` is a popover flag read from the live catalog.

Sites: apps/desktop/src/lib/desktop-slash-commands.ts::NO_DESKTOP_SURFACE
(46 rows) -> hermes_cli/commands.py::desktop_surface_registry (41 rows via
the dump) + TS_ONLY_NO_DESKTOP_SURFACE (5 rows). apps/desktop/src/AGENTS.md
updated.
2026-09-13 06:50:57 -07:00
teknium1
7fbe16aed5 test: exempt evals/ from the managed-runtime which() guard; make the OS-marker lister tolerate vanishing dirs
Two CI reds from the previous commits.

benchmark_browser_eval.py moved from scripts/ (exempt) to evals/ and its
bare shutil.which("npx") tripped test_no_unreviewed_bare_managed_runtime_
lookups. evals/ are standalone user-invoked benchmark programs, the same
class as scripts/ and skills/, and three other eval files already call
which() the same way; the guard just never saw them because the harness
that carried them lived under scripts/. Exempt the directory.

list_os_marked_tests.py used Path.rglob, which raises FileNotFoundError
when a __pycache__ directory disappears mid-scan (a sibling job in the
same workspace). The managed-runtime guard already switched to os.walk
for the identical TOCTOU; do the same here. Same output, sorted.
2026-09-13 06:06:46 -07:00
teknium1
b980495847 refactor: move live-model benchmark harnesses from scripts/ into evals/
scripts/ is for repo tooling (tests runner, release, installers, CI
checks). The tool-search live tests, the toolperf A/B eval and the browser
eval benchmark are offline benchmarks that spend model budget, which is
exactly what evals/ holds; evals/browser_use already cited
scripts/toolperf_abeval as "the same pattern".

- scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README
  -> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for
  the extra directory level; gitignore now covers evals/tool_search/out*/
- scripts/toolperf_abeval/ -> evals/toolperf_abeval/
- scripts/benchmark_browser_eval.py -> evals/browser_use/
2026-09-13 06:06:46 -07:00
teknium1
956967fbdc fix(ci): apps/ contract JSON edits run the Python lane
_PY_SKIP includes apps/, so a change to apps/shared/src/gateway-events.json
(or apps/desktop/src/lib/desktop-slash-registry.json) alone classified as
python:false and tests/tui_gateway/test_gateway_event_contract.py — the
only side that checks emitters against the JSON — never ran. "Either side
drifting turns both red" was true locally and false in CI for that
direction. A table of cross-language contract files now forces python:true
alongside frontend; two classifier rows pin it (they fail without the
allowlist).
2026-09-13 05:42:31 -07:00
teknium1
b2688440a9 fix(docker): rebootstrap re-seed temp is randomly named so a stale temp never blocks recovery
reseed_if_terminal created its temp as <auth>.rebootstrap.<pid>.tmp with
O_CREAT|O_EXCL. Boot-hook PIDs inside a container are near-deterministic,
so a run SIGKILL'd between create and replace leaves a same-named file and
every later boot hits FileExistsError - which main() swallows as
"error (ignored)", leaving the terminal-session recovery path dead until
someone deletes the temp by hand.

tempfile.mkstemp in the auth dir gives a random name at 0600 (stdlib only,
matching the script's no-hermes-imports rule); the fsync + os.replace +
unlink-on-failure semantics are unchanged.
2026-09-13 05:07:11 -07:00
teknium1
2be8e6147a refactor(secrets): every private-credential file is written by utils.atomic_json_write(mode=0o600)
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.

utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.

Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
2026-09-13 05:07:11 -07:00
teknium1
f7990661a0 fix(install): reuse probe skips the active venv; venv creation failure is fatal
Review finding: with the install's own venv activated (a re-run), `uv python
find <range>` returned venv/bin/python3, which setup_venv then deleted before
handing the dead path to `uv venv`; the stage still exited 0 and printed
"Virtual environment ready". Probe with --system so only base interpreters
qualify, and fail the stage when the venv was not created.
2026-09-12 08:32:09 -07:00
teknium1
5bdb4742b0 fix(install): reuse an installed supported Python instead of downloading 3.11
check_python only asked uv for 3.11 and, missing that, downloaded it —
a hard failure on hosts that cannot reach GitHub releases even when a
3.12/3.13 inside requires-python is already installed. Ask uv for the
supported range second and pin the venv onto that interpreter.

Same direction as #10824 (@gnanirahulnutakki), reduced to the shell
path; the PowerShell installer already has this fallback.

Refs #10778
2026-09-12 08:32:09 -07:00
teknium1
a7cf8ef8b2 fix(install): pick the .tar.gz Node.js archive when the host has no xz
`tar xf node-*.tar.xz` shells out to the xz binary; minimal Debian, DietPi
and WSL images ship tar without it, so extraction died mid-way and the
installer then failed on a missing directory. Select .tar.xz only when
`xz` is on PATH, in both the installer and the runtime node bootstrap.

Same approach as the earlier #4229 (@JoshuaMart) and #39541 (@karnull);
#11278 (@vominh1919) attempted an apt-only install of xz-utils instead.

Refs #11197
2026-09-12 08:26:17 -07:00
teknium1
a5228204b4 fix(branding): Caduceus in the WhatsApp bridge, installers and favicon too
Review finding on #109142: the JS bridge's DEFAULT_REPLY_PREFIX (the sender
actually used at runtime), scripts/install.sh, setup-hermes.sh and the site
favicon still carried the old glyph. Python and JS defaults now agree.
2026-09-12 08:25:54 -07:00
teknium1
e7794124da fix(skills-index): bound ClawHub owner enrichment so the scheduled index build finishes
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute
job timeout, so the live skills-index.json has been frozen at that date and every
`hermes skills search` fell through to live GitHub API calls (~500 inspect requests per
cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been
appending to #66616 four times a day since.

Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to
fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k
skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it.

- enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder
  without an owner (the link is a nicety; the index is not).
- build_skills_index.py passes an 8-minute budget.
- skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path
  (clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min).
2026-09-12 07:54:04 -07:00
MatthewHines
31bf8e504c fix(update): address review — neutral gate comment, pin empty-target skew case
- Rewrite the linux_gate comment to describe the environment fact
  (symlinked /home, kernel-canonicalised /proc/<pid>/exe) without
  local-patch markers or the unfiled-issue reference.
- Add test_empty_relaunch_target_falls_to_skew pinning the [ -n ... ]
  guard behavior so it cannot be simplified away.
- Add the missing trailing newline to the test file.
2026-09-12 05:09:59 -07:00
MatthewHines
d3b090a397 fix(update): canonicalise relaunch-target paths before the linux skew gate
The linux relaunch gate compares the running desktop's exe path
(relaunch target, read from /proc/<pid>/exe — kernel-canonicalised)
against the checkout's unpacked-app prefix with a raw case-pattern.
On hosts where /home is a symlink to /var/home (e.g. Fedora), the two
sides spell the same tree differently (/home/... vs /var/home/...) and
the gate false-positives "skew", telling the user to reinstall the
desktop app after every successful self-update.

Canonicalise both sides with readlink -m (which resolves existing
leading components without requiring the full path to exist, unlike
-f) before the prefix compare. No-op when both sides already agree.
2026-09-12 05:09:59 -07:00
Siddharth Balyan
b7b35a84b7 docs: remove the Nous guest free-tier guide (#108991) 2026-09-12 10:06:04 +00:00
Teknium
12d03eafed test(whatsapp): trim salvaged bridge tests to the quoted-media invariants (drop cache unit test) 2026-09-11 19:04:38 -07:00
ishangodawatta
e948ea8347 fix(whatsapp): don't drop bare quote-replies with resolved media
The empty-message guard only checked the reply's own body/hasMedia,
so a caption-less quote of a cached image (no text, no media of its
own) was dropped even though extractBridgeEvent had already resolved
quotedMediaUrls for it.
2026-09-11 19:04:38 -07:00
ishangodawatta
cd89e4b4a6 fix(whatsapp): resolve original media for quoted-media replies
Baileys' contextInfo.quotedMessage only ever carries a thumbnail-sized
stub for media, or nothing at all for an uncaptioned attachment — never
a way to fetch the original file. When a user replies to an earlier
photo/video/document/voice note with no caption on it (e.g. "did you
save this?" quoting an uncaptioned wedding invite image), the agent
saw no text and no media reference at all: it looked like the message
never had an attachment.

Add createQuotedMediaCache, a bounded in-memory cache (keyed by
chatId:messageId) of each inbound message's already-downloaded media
and text, populated as extractBridgeEvent processes every message.
When a later message quotes one of these, extractBridgeEvent resolves
quotedMediaUrls/quotedMediaType from the cache and falls back to a
human-readable quotedText ("sent an image", etc.) when the quote had
no caption to extract. The adapter folds resolved quoted media into
the event's own media_urls/media_types — reusing the existing
vision/audio pipeline and the existing _is_allowed_bridge_path path
validation — rather than adding a parallel reply-media code path.

Reimplements the same feature as #52875 (credit: dhruvkej9) against
current main, whose 11627fdcb refactor (native polls, locations, rich
inbound metadata) moved this code into bridge_helpers.js and made that
PR's diff no longer apply cleanly.
2026-09-11 19:04:38 -07:00
Teknium
0dcadf6f41 revert: remove Collective Wisdom V1 (#94266)
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three
model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces.

Later non-Wisdom work on shared files (guest onboarding i18n, dashboard
startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call
sites and config were stripped from those files.
2026-09-11 11:54:49 -07:00
KoNit-K
0b8daf30aa fix(bootstrap-installer): stamp setup app version from release semver
Fixes #107177

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-11 16:10:47 +02:00
shannonsands
a6ee31f55a feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation

* feat(wisdom): add private contribution loop

* feat(wisdom): add managed consumption workflows

* fix(wisdom): close cross-repository safety gaps

* fix(wisdom): align local package and lifecycle policy

* fix(wisdom): require explicit profile setup

* docs(wisdom): repin reconciled gateway head

* fix(wisdom): fence content downloads and approval receipts

* docs(wisdom): record generation-fenced downloads

* docs(wisdom): record unified delivery PR

* fix(ci): stop passing invalid classifier inputs

* docs(wisdom): remove internal requirements ledger

* feat(wisdom): localize dashboard and desktop copy

* feat(wisdom): complete local contribution and consumption UX

* style(wisdom): satisfy desktop lint

* chore(wisdom): refresh requirements pin

* test(dashboard): allow formatted profile copy

* test(wisdom): stabilize desktop interaction coverage

* fix(wisdom): surface dashboard action failures

* fix(wisdom): add repeatable Portal demo login

* feat(wisdom): add actionable skill notifications

* feat(wisdom): add notification install and update actions

* fix(wisdom): make Telegram skill alerts actionable

* fix(wisdom): always refresh demo Agent login

* feat(wisdom): embed Telegram notification actions

* fix(wisdom): preserve Telegram notifications after actions

* fix(wisdom): keep Telegram notification cards readable

* feat(wisdom): add Telegram candidate approval flow

* feat(wisdom): explain Telegram qualification reasons

* fix(wisdom): reconcile cross-surface candidate actions

* feat(telegram): add Collective Wisdom management command

* chore(wisdom): refresh Gateway contract pin

* chore(wisdom): advance Gateway contract pin

* feat(wisdom): align command UX across clients

* feat(slack): add Collective Wisdom management parity

* feat(wisdom): add security and professionalism reviews

* feat(wisdom): add first-time qualification guidance

* feat(wisdom): simplify qualification sharing choices

* feat(skills): add optional editorial metadata

* feat(wisdom): enrich legacy skill presentation

* fix(wisdom): harden review and update boundaries

* fix(wisdom): emit canonical review timestamps

* fix(wisdom): align with merged gateway and main

* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)

- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
  7-day evidence builder that excludes bundled/hub/managed skills and
  dismissed/handled/recently-suggested content hashes, strict pydantic
  schemas for agent output with repair-or-reject, fixed copy templates
  (Share / Teammate / Published / Update / Mute), idempotent retried
  delivery ledger with stale-action resolution, weekly review job,
  resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.

* wisdom: agent-led renderers and button action dispatcher

- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
  is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
  ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
  packaging flow, Install/Update -> plan command. Never publishes/installs.

* wisdom: CLI verbs, agent_led config default, conversational catalog skill

- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
  verbs, share/install flows and fixed notification templates.

* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons

- gateway housekeeping tick calls maybe_run_weekly_review with a home
  channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
  duration keyboard, send_wisdom_agent_recommendation rich card + fallback.

* fix(wisdom): integrate local mediation and harden model and setup boundaries

* fix(wisdom): honor authoritative recommendation policy and defer on failure

* fix(wisdom): synchronize opaque suppression and recheck delivery preferences

* feat(wisdom): route weekly selection through the session-owned assessment queue

* fix(wisdom): prepare and submit the reviewed generated share package

* feat(wisdom): separate native Share preparation from publication consent

* feat(wisdom): sync native mute choices through a leased preference outbox

* feat(wisdom): bind native mute controls to durable preference choices

* feat(wisdom): add scoped desktop and dashboard notification settings

* fix(wisdom): revalidate feed recommendations before assessment and delivery

* fix(wisdom): persist validated delivery receipts before completing notices

* feat(wisdom): add private notification claim and receipt client

* Persist Wisdom send reservations and recover delivery acknowledgements

* Route legacy Wisdom controls through current native review

* Add typed private Wisdom operation outcome client

* fix(wisdom): make agent-led advice usable in the local demo

* fix(wisdom): keep requested consent outside proactive limits

* fix(wisdom): distinguish unavailable assessments and preserve digest text

* fix(wisdom): assess ongoing usefulness beyond the current task

* fix(wisdom): restore immediate qualification sharing controls

* fix(wisdom): separate qualification review from installation advice

* fix(wisdom): collapse review checklists and simplify sharing copy

* fix(wisdom): show compact sharing progress and publication receipts

* fix(wisdom): require credential prefixes rather than matching skill names

* fix(wisdom): finish package checks before presenting sharing consent

* fix(wisdom): scan local skills before qualification cards

* fix(wisdom): update moderation results on existing sharing cards

* fix(wisdom): keep sharing review accessible from receipt cards

* fix(wisdom): align mediated review cards and collapsible checks

* fix(wisdom): clarify clean security summary wording

* fix(wisdom): normalize consent plans and add explicit recheck

* fix(wisdom): keep install and update receipts concise

* fix(wisdom): collapse assessments and deduplicate operation cards

* fix(wisdom): restore private Portal review from native cards

* fix(wisdom): sync Portal publication to original consent card

* fix(wisdom): show local skill version on sharing cards

* fix(wisdom): skip agent recommendations for self-published versions

* fix(wisdom): simplify candidate notices and local-edit recovery copy

* feat(wisdom): submit locally reviewed packages with one confirmation

* feat(wisdom): expose safe receipt and outcome sync recovery

* wisdom: onboarding notice says detect and share, names the user's own skill

Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark

Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.

* wisdom: one opener, no approval line, ask to share after the skill is shown

Product owner review of the candidate card.

- The Hermes written card now opens with the same sentence as the fixed card
  ("Your organisation has enabled Collective Wisdom, a feature designed to
  automatically detect and share useful skills across all team members.")
  instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
  and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
  It is now the last line, after the skill name, description, why suggested
  and the checks, and reads "Would you like to share it?" (matching the
  agent led template wording).

Tests updated for the new order; proposalNotice removed from all desktop locales.

* wisdom: American spelling, organization

Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.

* wisdom: candidate card copy round 4 (owner review)

Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:

1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
   the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
   card (Telegram rich card and plain fallback, legacy agent-led share
   template).
3. The skill name and description are labelled: "Skill name: <name>" and
   "What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
   inappropriate content found)" with no per-check bullets and no "Pass";
   a failed review reads "Needs a look before sharing at work (possible
   inappropriate content)" and lists only the checks that flagged
   something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
   details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
   days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
   editorial_name, a simple one_line_description and a compelling
   why_coworkers_benefit under 300 characters; "Be concise and
   convincing." becomes "Be concise and compelling: the goal is that the
   user wants to share it."

Tests updated for the new strings; review_text() gains direct coverage.

* wisdom: re-apply owner copy after rebase

- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice

* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors

Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.

* fix(wisdom): reconcile optional SDK tests and frontend lint

* fix(wisdom): default to agent-written notification summaries

* fix(wisdom): restore deferred install review and browse controls

* feat(wisdom): inspect installed setup with exact package provenance

* feat(wisdom): run native-approved installed setup steps with durable evidence

* fix(wisdom): recover interrupted setup with explicit native consent

* feat(wisdom): hand native installs into guided setup review

* fix(wisdom): continue requested setup with fixed notification copy

* fix(wisdom): preserve setup while waiting for a session model

* fix(wisdom): expose canonical setup review controls on desktop

* fix(wisdom): resume setup after recorded automatic updates

* fix(wisdom): make missing setup prerequisites recheckable

* chore(wisdom): align Agent with verified Gateway contract

* fix(wisdom): stop guessing team slugs in portal links

* fix(wisdom): retire pending advice on account sign-out

* fix(wisdom): cancel advice after terminal account revocation

* fix(wisdom): fence feed responses across account sign-out

* fix(wisdom): checkpoint signed-out feed before reactivation

* fix(wisdom): link proactive advice to scoped notification settings

* fix(wisdom): coalesce queued publication recommendations by version

* fix(wisdom): keep package review navigation local and deferable

* fix(wisdom): reflect installed state in discovery controls

* fix(wisdom): show exact checks before command confirmation

* chore(wisdom): pin bounded analytics privacy contract

* chore(wisdom): pin retired legacy notification contract

* feat(wisdom): review publisher usage with exact sharing copy

* fix(wisdom): align discovery and review check summaries

* fix(wisdom): show expired consent before confirmation

* fix(wisdom): require fresh review for legacy install controls

* fix(wisdom): preserve review expiry across check toggles

* fix(wisdom): retain update policy in native install reviews

* fix(wisdom): surface failed native card edits

* fix(wisdom): persist local command approval reviews

* fix(wisdom): use saved approvals for messaging commands

* test(wisdom): provide scan result in setup handoff fixture

* test(wisdom): exercise Telegram approvals with saved review state

* fix(wisdom): retain suppression policy for offline deferral

* fix(wisdom): reconsider candidates after deferred suppression expires

* fix(wisdom): bind review checks and report verified readiness separately

* fix(wisdom): persist accepted publication intent and recover exact outcomes

* fix(sync): pin UTF-8 tree ordering across writers

* chore(wisdom): pin organisation-scoped Gateway authorization

* fix(wisdom): restrict consent delivery to user-facing sessions

* chore(wisdom): refresh reviewed Gateway contract pin

* fix(wisdom): preserve kept tools in Blank Slate exclusions

* test(auth): reset anonymous fixture with a profile-scoped cache

* fix(wisdom): gate local surfaces and work on current profile entitlement

* fix(wisdom): invalidate quiet tool cache on entitlement changes

* test(wisdom): authorize local consent gateway fixtures

* fix(wisdom): keep entitlement decoding free of native crypto imports

* test(wisdom): provide local entitlement to demo CLI subprocess

* ci: leave upstream workflow unchanged in Wisdom PR

* fix(wisdom): ship package and contracts in Nix wheels

---------

Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
2026-09-11 19:04:06 +10:00
Teknium
8068c09432 fix(desktop): verify the Windows update receipt against the checkout, not cwd
The Windows hand-off script runs from the PRE-update checkout, spawned from
HERMES_HOME by the Desktop. e7eaff68 shipped the post-update verify step
without the cwd pin that landed a day later, so verify_windows_desktop_update(Path.cwd())
looked for apps/desktop/release under ~/.hermes and reported a healthy,
fully updated install as "The updated Desktop executable is missing" (exit 8,
error dialog, app relaunched fine).

Derive the root from the imported hermes_cli package instead; the ps1 no
longer passes a cwd. The cwd pin stays as belt for the other child steps.
2026-09-10 01:25:45 -07:00
Teknium
e74c4a00ca Merge pull request #69446 from NousResearch/feat/plugin-catalog
feat: plugin catalog — curated SHA-pinned plugin index (CLI, admission CI, docs, dashboard)
2026-09-09 09:22:21 -07:00
Teknium
e8bac35a40 refactor(whatsapp): bridge.js reuses the helpers' getMessageContent
bridge.js carried its own copy of getMessageContent with the single-layer
envelope list, alongside an unused getContextInfo. With envelope peeling now
living in bridge_helpers.js, the copy would drift (a nested envelope around
a pollUpdateMessage is peeled by the helpers but not by the copy), so import
the shared one instead of keeping a second unwrapping list.
2026-09-09 09:21:12 -07:00
Konstantin Khlopkov
7ecd4d14a5 fix(whatsapp): peel nested envelopes and restore immediate payload return
A quoted message can carry its own envelope stack (ephemeral
wrapping viewOnce wrapping the payload), and a reply itself can be
enveloped: both layers now unwrap iteratively (bounded) before quote
extraction. Peeled top-level envelopes return their inner message
immediately, as before — the template/buttons/list branches only
apply to unenveloped messages.
2026-09-09 09:21:12 -07:00
Konstantin Khlopkov
5de769956f fix(whatsapp): unwrap quoted-message envelopes before extracting quote text 2026-09-09 09:21:12 -07:00
Teknium
d47adec28f Merge origin/main into feat/plugin-catalog
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
2026-09-09 04:15:27 -07:00
Teknium
bf53ff00a7 fix(config): one bounded backups/config/ dir replaces four config.yaml.bak schemes
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.

hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.

Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
2026-09-09 02:36:00 -07:00
Teknium
c32e0acb0e test(desktop): exercise production cwd setup in Windows self-test 2026-09-08 17:13:48 -07:00
fangliquanflq
610b6731d4 fix(desktop): pin Windows update handoff cwd 2026-09-08 17:13:48 -07:00
ethernet
c8aa5608c2 Merge pull request #101420 from ethernet8023/ethie/desktop-update-tests
test(install): cross-OS install/update E2E matrix (windows, macos, linux)
2026-09-08 10:30:44 -04:00
Teknium
a662f9d513 fix(desktop-update): discard failed profile allocation output 2026-09-07 06:09:56 -07:00
Teknium
ca812ba3b5 fix(desktop-update): atomically claim the temporary browser profile
Use mktemp -d before launching the optional UI; skip UI if allocation fails. Native Chrome collision and allocation-failure probes preserve preexisting directories.
2026-09-07 06:09:56 -07:00
Teknium
1bd7364d8e fix(desktop-update): clean only the captured shim profile
Track the path actually launched, preserving the no-UI case and unrelated profiles. Adapted the ownership approach from #104362.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 06:09:56 -07:00
Rohith Pariki
97f74b8361 fix(desktop-update): clean up throwaway browser profile directory
Deletes the temporary --user-data-dir used by the update UI shim when the browser process is shut down, preventing ~100MB leaks per update. Fixes issue #104350.
2026-09-07 06:09:56 -07:00
yoniebans
15860dfe9f Merge ethie's suite hardening; her input preparation supersedes the dispatchEvent fallback
Both sides fixed the swallowed-click class on the onboarding picker.
Hers is the root cause: persistent 100% zoom through the app's own
setting, verified from both the renderer IPC and the BrowserWindow, and
re-applied before the dismiss loop — scale drift is what moved real
click points onto the wrapping container. The dispatchEvent fallback is
dropped: it bypassed hit-testing, so a leg could pass where a real
user's click would fail.

Comment-only conflicts in managed_uv.py and main_install_repair.py
resolved by keeping the fuller mechanism text (import-order reach and
the legacy hand-off scope).
2026-09-07 14:58:54 +02:00
Teknium
5ce8e974c2 fix(desktop): reject incomplete Windows builds before success receipts
Native Windows run 34096838164 reports false success for absent and corrupt executables, missing bundle files, missing chunks and missing or stale stamps. Reuse the existing build identity and PE validators, and check interpreter presence before waiting for Desktop. Preserve dependency recovery and exit-2 refusal behavior.

Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-07 05:55:26 -07:00
Teknium
90ac288c7d fix(desktop): verify updated runtime before success receipt
Port the runtime verification portion of #104692 after native run 34095483533 reproduced ok=true for a zero-exit controlled child that removed its runtime module. Artifact/build-stamp validation remains unaddressed.

Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-07 05:55:26 -07:00
Teknium
faf5b42a82 fix(desktop): fail missing Windows updater handoffs
Salvage the missing-target guard from #104692. Native Windows run 34094671567 returned exit zero for the absent maintained script. Full runtime/artifact completion remains separate.

Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-07 05:55:26 -07:00