Commit Graph

351 Commits

Author SHA1 Message Date
teknium1
1536e75bfe docs(website): generator prunes stale skill pages; CI fails when the committed docs drift
`generate-skill-docs.py` now deletes every page under `user-guide/skills/{bundled,optional}`
it did not write this run, together with the zh-Hans mirror twin, so a skill that moves,
merges or leaves the shipped set takes its page with it instead of lingering as an orphan
that cross-links still reach (24 such pages after the shipped-set slim, plus 21 zh-Hans
copies whose English page was already gone).

The Docs Site Checks workflow regenerated the docs and never compared the result with the
committed copies, which is what GitHub renders; it now fails with a pointer to the generator
when they differ. Also regenerates the two pages that drifted since the salvaged commits.
2026-09-20 12:54:07 -07:00
teknium1
3999096d18 ci: forbid literal /tmp paths outside a burn-down baseline
scripts/check_no_tmp_literals.py flags /tmp path tokens in production code, skills,
docs and prompt strings (tests, CI workflows, Dockerfiles, lockfiles, i18n mirror,
code comments and docstrings exempt; ${TMPDIR:-/tmp} idiom exempt). Opt out one line
with 'no-tmp: ok — <why>' on the line or the line above. _BASELINE lists pre-existing
hits per file: growth fails, burn-down is advisory (--strict-baseline / --print-baseline
to refresh). Wired into lint.yml next to check_compat_pointers.
2026-09-19 10:44:26 -07:00
teknium1
d0dbf2cbb6 fix(install): resolve uv shims before salvage and validate the copy in place
Install-Uv accepted any file at $HermesHome\bin\uv.exe, and copied whatever
`Get-Command uv` returned into that location. Chocolatey's bin\uv.exe is a
ShimGen launcher that locates ..\lib\uv\tools\uv.exe RELATIVE to itself, so
the copy is dead on arrival; `& exe --version` does not throw on a nonzero
exit, so the launcher passed the try/catch and the Python stage then failed
with "Python 3.11 not available" (#110350). The re-run path trusted the same
broken copy again.

Building on KoNit-K's Test-ManagedUvBinary and its three call sites:

- Test-ManagedUvBinary merges stderr, relaxes the error preference, and
  returns the `uv <version>` line only on exit 0 -- a launcher's error text
  can no longer surface as "Managed uv found (Cannot find file ...)".
- Resolve-UvShimTarget maps a candidate to the standalone binary before the
  copy: `<name>.shim` sidecar (Scoop), the Chocolatey bin\ -> lib\<pkg>\tools\
  layout, symlinks (winget Links\); other reparse points (WindowsApps
  app-execution aliases) have no copyable file and skip the salvage.
- The salvage rung validates the candidate where it lives, copies, then
  validates the COPY at its new location and removes it on failure, so the
  stage fails honestly instead of reporting success over a dead launcher.
- scripts/tests/test-install-ps1-uv-shim-validation.ps1 drives the real
  Install-Uv with compiled fake uv binaries (a working uv and a
  location-relative launcher) under stubbed installer rungs; wired into
  installer-tests.yml for pwsh 7 and Windows PowerShell 5.1.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-19 02:44:02 -07:00
Siddharth Balyan
3752f89ff8 fix(ci): a failing workspace check keeps its output in the log (#115214)
`run-workspace-checks.mjs` buffers each check's output, writes it in one
piece, and then called `process.exit(1)`. On a pipe, stdout is asynchronous,
so the exit dropped what was still buffered: the log of a failing check
stopped in the middle, before the test runner's failure summary.

Set `process.exitCode` and return, so the process ends after stdout drains.
2026-09-19 02:17:14 -07:00
teknium1
36fb2be926 fix(website): cross-page doc links resolve on GitHub as well as on the site
Docs under website/docs were linked with Docusaurus site routes
(`](/getting-started/installation)`, `](/docs/user-guide/x)`), which GitHub's
file viewer resolves as repository paths and 404s (#114428). Relative Markdown
file links (`](../user-guide/x.md#anchor)`) are followed by both GitHub and
Docusaurus, so that becomes the authoring convention:

- website/scripts/check_doc_links.py lints hand-authored EN + zh-Hans pages
  for route-style links (`--fix` rewrites them, refusing any route that maps
  to no doc file); wired into the Docs Site Checks workflow and
  tests/website/test_check_doc_links.py.
- generate-skill-docs.py emits the same relative form for related-skill and
  catalog links instead of `/docs/user-guide/skills/...`.
- src/remark/relativeDocLinks.js rewrites `./x.md`/`../x.md` to
  content-root-absolute `/x.md` before Docusaurus resolves links, so a
  relative link still resolves in the zh-Hans build when source and target
  sit on different sides of the translation fallback (Docusaurus resolves
  `./`/`../` only against the source file's own directory).
- website/README.md states the convention and points at the checker.
2026-09-18 14:27:04 -07:00
teknium1
73f39086f6 fix(skills-index): retry only on a failed fetch; give the scheduled build time for slow ClawHub pages
Trim the salvaged retry to the case the logs show: `_get_json` returning None
(timeout / non-200) is a hole in the cursor walk, not the end of the catalog.
A dict page with an empty item list stays terminal as before, and the
interactive browse path (12 s budget) no longer sleeps up to 8 s per retry.
Constant `CATALOG_PAGE_RETRIES` replaces the local literal.

Workflow: `skills-index.yml` `timeout-minutes` 50 -> 120. Since Sep 17 ClawHub
serves ~10 s per 200-item page (was ~2.5 s), so the ~400-page walk alone takes
60-70 min; the 06:28 UTC runs on Sep 17 and Sep 18 were cancelled at 50 min with
the clawhub crawl still running, and the 18:19 run stopped at 12,145 skills
(< 20,000 floor) after one failed page. Both halves are needed for the index to
ship again.
2026-09-18 12:51:49 -07:00
Robin Fernandes
07a05437b3 fix(desktop): support Privy and NAS portal sessions in Cloud sign-in 2026-09-18 00:00:23 -07:00
teknium1
f253f25b78 feat(plugins validate): --install-deps installs the declaration before the capability probe
Catalog CI validates each pinned tree in a venv that has only hermes-agent, so any plugin whose
code arrives through pyproject dependencies (the wrapper shape from #113851) failed the probe with
"No module named ...". The flag runs the same constrained install `plugins install` would, then
probes; the workflow passes it. Live: mnemosyne pin fails without the flag, passes with it.
2026-09-17 11:21:42 -07:00
teknium1
290bdc3c76 fix(update): let the Windows repoint run; refuse a minor-line jump
The self-lock/holder preflight (#99711) deferred the repair on the theory
that the updater's own mapped python.exe makes the venv rename impossible.
Live on windows-latest, a process executing from venv\Scripts\python.exe
(and the repo's real .venv with cp311 .pyd extensions loaded) does NOT block
the rename; what blocks it with WinError 5 is any ordinary handle under the
tree: a process cwd, an open file, a sync client. Since sys.executable is
always under the venv on Windows, the deferral fired on every run and no
Windows install could repair from `hermes update`. Retire it and its tests;
the pyvenv.cfg repoint needs no rename and has none of that exposure. The
wine2e lane now runs the cutover probes instead.

The repoint keeps the live venv's site-packages, so provisioning's
fall-forward to the next minor (3.11 -> 3.12, #76106) must not be pointed at
a cp311 tree: refuse before touching pyvenv.cfg. The live Windows test holds
the venv the way the field does (a child with cwd inside it), asserts the
rename path fails (the symptom) and the repoint succeeds; the rollback and
minor-guard tests run on every host.
2026-09-17 08:44:33 -07:00
teknium1
bf1c28b480 ci: fail lint when a test fakes macOS without @pytest.mark.macos_only (#111866)
scripts/ci/check_os_marker_fakes.py flags test files that make the
interpreter believe it is on macOS (is_macos -> True, sys.platform ->
"darwin", platform.system -> "Darwin") while carrying no `macos_only`
marker: the macOS lane imports only marked files, so such a file is green
on Linux over a faked branch and never runs on the host it exists for.

A `# os-marker: ok — <why>` comment opts a host-independent line out.
_BASELINE holds the files that already faked macOS when the check landed;
a stale entry fails the check so the list can only burn down. Wired into
lint.yml next to the compat-pointer check; two invariant tests cover a
flagged fake vs marked/opted-out files and host-honest platform reads.
2026-09-15 21:49:33 -07:00
teknium1
1cfa892db1 fix(desktop): make the zsh probe test legs visible and run them on CI
The zsh login-shell legs in remote-lifecycle.test.ts and
ssh-connection.test.ts silently returned when zsh was missing, and the
js-tests runner image ships no zsh, so the #111949 coverage never ran on
CI and a wrapper regression stayed green.

- js-tests.yml: install zsh on the Linux runner before the checks.
- Both legs now report vitest skips ('zsh not installed') instead of
  passing; the ssh-connection leg is its own test so the skip is visible.
- Docs: note the zsh degraded mode (no process-group kill for a hung
  probe's grandchildren) in the SSH connection guide.
2026-09-15 18:46:32 -07:00
teknium1
a4f8f91e50 ci: run the OSV lockfile scan weekly against main, not on every PR
The scan is detection-only and its findings are the repo-wide baseline of
CVEs in pinned dependencies — identical for every PR, unrelated to any
PR's diff. Reporting that baseline in each PR's review comment read as
"this PR has 76 vulnerabilities" to contributors, and the SARIF upload
tripped GitHub's per-installation API rate limit during merge trains.

The scheduled weekly run (plus workflow_dispatch) keeps feeding the
Security tab; the per-PR workflow_call, the review_status wrapper job,
and the orchestrator's now-unneeded SARIF permissions are removed.
2026-09-15 11:43:45 -07:00
teknium1
f13a87e610 ci: advisory profile-scope pattern lint on the lines a PR adds
scripts/check_profile_scope_patterns.py runs the validated hazard regexes in
scripts/ci/profile_scope_patterns.json (18 of the 31 campaign patterns: every one has a scope_hint
and hits <= 50 sites on main; the wider ones are review greps, not lint) against the lines added
vs the PR base and prints file:line, pattern id/class and why. Always exits 0: most shapes have
legitimate sites (a standalone `hermes -p x` process where environ IS the profile), so the
reviewer reads each finding against its scope hint. Wired into lint.yml beside the public-surface
diff with continue-on-error.

Proof: the pre-fix tools/bot_relay.py (`env = dict(os.environ)`, before the served_profile_child_env
change) is flagged as P05/C2; the fixed file and this branch's diff vs main report 0 findings.
Test: a fixture with the hazard is flagged on the right lines, the scoped-builder version is not,
and the line filter hides hits outside the added range.
2026-09-15 10:59:22 -07:00
teknium1
471d8a427a ci(plugin-catalog): probe stars only from the scheduled skills-index run, in one GraphQL request
Aligns the star ranking with the rule the skills index already follows:
GitHub is consulted only by the twice-daily skills-index.yml schedule, whose
artifact every docs deploy reuses. deploy-site.yml now runs
fetch-plugin-stars.py without --probe (reuse-only: artifact → live site copy →
disk → empty) and cannot call the API at all, so a same-day merge train adds
zero requests regardless of cache age.

The probe itself collapses from one REST call per repo (~26 today, growing
with the catalog) to a single GraphQL query with aliased repository fields,
so the scheduled run costs one request no matter how big the catalog gets. A
failed probe (rate limit, renamed repo, bad token) keeps the previous counts.
2026-09-15 04:39:38 -07:00
teknium1
d228013832 fix(ci): feed the timeout scaler only healthy durations, and wire the cache in CI
Greptile's two findings on the original PR were both right.

1. The scaler read test_durations.json from the checkout, but CI ran on
   a fresh runner where that file never exists (it is gitignored and the
   slicing-era artifact/merge job that produced it is gone). The feature
   was inert exactly where the false FLAKY kills happen. tests.yml now
   restores the most recent main-saved cache before the run (PRs read
   only) and saves it after a green push to main, mirroring the
   ci-timings-baseline restore/save pattern already in ci.yaml.

2. _save_durations persisted every file's total subprocess wall,
   including the ~cap of a timed-out attempt and the retry-summed wall
   of a FLAKY file. With the scaler that compounds: a hang cached at
   ~300s earns 900s next run, then ~900s cached earns 2700s, until the
   job timeout is the only bound. _clean_pass_durations drops failed and
   FLAKY files from the write so a file's cached duration is always a
   first-attempt-clean measurement; those files keep their previous
   known-good entry.

Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
2026-09-15 03:47:55 -07:00
teknium1
3e2e2c50eb feat(plugin-catalog): rank entries by GitHub stars, probed at most once a day
Catalog entries sort official → stars desc → name, both in browse shelves and
filtered grids, with a ★ pill on each card linking to the repo's stargazers.

Rate-limit discipline is the design constraint: the docs site deploys many
times a day and shares one GitHub App API budget with every other workflow
(tonight's merge train got rate-limited on unrelated uploads). So
website/scripts/fetch-plugin-stars.py first fetches the live site's own
plugin-stars.json (a CDN GET, not the API); if that cache is under 24h old it
is reused verbatim and GitHub is never called. Only a stale cache triggers one
GET /repos/{owner}/{repo} per unique catalog repo, and a 403/429 mid-run keeps
the previous counts instead of zeroing them. extract-plugins.py merges the
cache into plugins.json (`stars`) and plugins-meta.json (`starsFetchedAt`), and
the page footnote says when the ranking was last refreshed.
2026-09-14 21:24:33 -07:00
teknium1
5db6d40874 ci: stop the advisory OSV scan from gating merges
osv-scanner.yml documents itself as detection-only (fail-on-vuln: false,
findings land in the Security tab) yet all-checks-pass listed it in needs,
so any failure result blocked the merge. In practice the failures are not
vulnerabilities: the "Upload to code-scanning" step hits GitHub's
per-installation API rate limit whenever several PRs run at once, and a
merge train of catalog entries went red on it across the board. The scan
still runs on every PR and weekly on main; it just reports instead of gating.
2026-09-14 19:15:00 -07:00
teknium1
48763a4d01 docs(plugin-catalog): replace the 2-week pin-maturity rule with a no-self-updater rule, enforced in CI
Teknium's ruling: catalog plugins do not need the 2-week maturity window,
but they may not ship an in-app updater that downloads and replaces their
own files, because that makes the reviewed SHA pin decorative. Rule 3 in
the README and item 5 on the docs page now say so, and plugin-catalog-ci
fails an entry whose catalog build both fetches from GitHub releases/raw
and writes or renames plugin files (either half alone is allowed).
2026-09-14 17:17:19 -07:00
Teknium
915f23efc7 Port from nearai/ironclaw#7756: bound the last unbounded CI job
IronClaw's #7756 swept every unbounded CI operation (apt hangs, uncapped
jobs, external downloads). Same sweep here found exactly one gap: the
osv-scanner emit-status wrapper job had no timeout-minutes, so a wedged
artifact download could hold a runner for GitHub's 6-hour default. Every
other job across all 30 workflows is already bounded. Capped at 10m.
2026-09-13 19:41:37 -07:00
teknium1
d10bb2ab6f test: make tests/ mirror the source tree; drop issue numbers from filenames
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.

Parallel directories for one source package, folded into the mirror:
  tests/acp        -> tests/acp_adapter   (its __init__/conftest move with it)
  tests/cli        -> tests/hermes_cli    (prompt_toolkit fixture merged into
                                           hermes_cli/conftest.py)
  tests/run_agent  -> tests/agent         (backoff fixture becomes
                                           agent/conftest.py)
  tests/relay      -> tests/gateway/relay
  tests/state      -> tests/hermes_state

246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.

Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.

Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).

Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
2026-09-13 09:18:02 -07:00
teknium1
e7794124da fix(skills-index): bound ClawHub owner enrichment so the scheduled index build finishes
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute
job timeout, so the live skills-index.json has been frozen at that date and every
`hermes skills search` fell through to live GitHub API calls (~500 inspect requests per
cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been
appending to #66616 four times a day since.

Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to
fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k
skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it.

- enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder
  without an owner (the link is a nicety; the index is not).
- build_skills_index.py passes an 8-minute budget.
- skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path
  (clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min).
2026-09-12 07:54:04 -07:00
Teknium
f7a973e3d7 ci(lint): diff the PR head ref so the public-surface deepen loop can actually find a merge-base
The advisory step ran `git merge-base origin/main HEAD` where HEAD is the
checked-out refs/pull/N/merge commit. GitHub recomputes that synthetic merge
commit whenever main moves, so by the time the job runs the checked-out SHA is
often no longer reachable from any ref. `git fetch --deepen` on an
unreachable commit can never pull in its parents, so all four fetches in the
loop (200 + 3 x 1000) ran to completion, burned 4m15s, and still ended in
"public-surface: no merge-base" (run 34285107265, job 102259026802). The
step-level timeout keeps that overrun from cancelling the blocking job; this
commit makes the step finish on the first fetch instead.

Fetch refs/pull/N/head explicitly (always reachable server-side) and diff
that ref against the base; the merge commit's own content is irrelevant to a
symbol diff.

Raise the step timeout from 2 to 3 minutes: one --deepen=200 fetch took ~56s
and a --deepen=1000 ~67s on the runner, and main gains ~170 commits a day, so a
branch a day old legitimately needs both. 45s of blocking steps + 3 min still
fits the 5-minute job budget.
2026-09-09 12:19:41 -07:00
PRATHAMESH75
981dc9c367 ci(lint): bound the advisory public-surface step so it can't time out the blocking footguns job (#106103)
The 'Public-surface diff vs base (advisory)' step already carries
continue-on-error: true, so its own failure never fails the Windows-footguns
job. But continue-on-error does not shield the job-level timeout-minutes: 5:
when the deepen/fetch loop runs long (a distant or missing merge-base), the
step eats the whole job budget and the job is cancelled at ~5m even though the
blocking footgun/compat checks passed and the public-surface check is advisory.

Give the advisory step its own timeout-minutes: 2. A step overrun is then
killed and, via continue-on-error, kept off the job outcome the same way a
step failure already is, so the blocking job can finish under its budget.
Fixes #106103.
2026-09-09 12:19:41 -07:00
Teknium
41b4555ed9 feat(plugins): catalog is the sole discovery system — re-port onto main's layout
- hermes_cli/plugins_cmd_catalog.py: new sibling owning resolution, the
  .hermes-catalog.json provenance sidecar, search/info/validate, re-pin on
  update, and the dashboard/TUI payload builders. plugins_cmd.py only
  gains the hooks (cmd_install catalog branch, cmd_update / dashboard
  update re-pin, dashboard_install_plugin catalog_name + kill list,
  dispatch entries); the community index (plugin_index.py) is gone.
- hermes_cli/plugin_catalog.py: catalog_dir parameter replaces the
  test-only HERMES_PLUGIN_CATALOG_DIR env var; live refresh reads ONE
  published document (/docs/api/plugin-catalog.json, 6h cache, in-tree
  fallback) instead of the unauthenticated GitHub contents API (60 req/h,
  1 request per entry); in-tree and live removals are unioned so a stale
  cache can never un-block.
- Catalog route lives in web_routers/dashboard_ui.py (the facade is off
  limits); _plugin_runtime_status shared from web_server_dashboard.py;
  hub rows carry removed_reason. TUI plugins.manage gains catalog_name
  install, catalog row fields and an update action.
- plugin_validate: the probe context honours ctx.get_config defaults
  (real plugins do int(ctx.get_config("timeout", 180)) in register()).
- Installed-state merge matches through the sidecar's catalog_name
  first — catalog names rarely equal manifest names.
- extract-plugins.py emits plugin-catalog.json; deploy-site triggers on
  plugin-catalog/** so entry merges republish it.
2026-09-09 04:38:01 -07:00
Teknium
d47adec28f Merge origin/main into feat/plugin-catalog
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
2026-09-09 04:15:27 -07:00
yoniebans
15860dfe9f Merge ethie's suite hardening; her input preparation supersedes the dispatchEvent fallback
Both sides fixed the swallowed-click class on the onboarding picker.
Hers is the root cause: persistent 100% zoom through the app's own
setting, verified from both the renderer IPC and the BrowserWindow, and
re-applied before the dismiss loop — scale drift is what moved real
click points onto the wrapping container. The dispatchEvent fallback is
dropped: it bypassed hit-testing, so a leg could pass where a real
user's click would fail.

Comment-only conflicts in managed_uv.py and main_install_repair.py
resolved by keeping the fuller mechanism text (import-order reach and
the legacy hand-off scope).
2026-09-07 14:58:54 +02:00
yoniebans
32ed2061bb Merge upstream main (fc8d15d779): freshen before PR push 2026-09-07 10:15:01 +02:00
ethernet
692ef3294c test(install-e2e): remove tracing after verified July handoff fix 2026-09-06 22:26:35 -04:00
ethernet
1c8b4d5faf fix(install-e2e): stage classified receipts at an artifact-safe path 2026-09-06 20:10:33 -04:00
ethernet
49bf392f0a feat(install-e2e): classify exact known failures with report footnotes
Keep unknown failures red, rotate evidence per attempt, and emit receipts for signature-confirmed historical cases. Add CI-only diagnostics and an exact-tag input for the unresolved July hand-off.
2026-09-06 19:46:10 -04:00
Teknium
857218e6da ci(lint): deepen the shallow checkout before the public-surface diff; the step is advisory and never fails the job
The round-2 change made check_public_surface refuse to report a clean diff
without a merge-base (exit 2). That correctly exposed that the lint job's
depth-1 checkout plus a depth-1 fetch of the base has NO merge-base, so the
advisory step had been silently reporting 0 drops on every PR. The step now
deepens both sides until a merge-base exists and carries continue-on-error so
an advisory check can never block the Windows-footguns job it rides in.
2026-09-06 13:27:48 -07:00
Teknium
8cbb2ce764 feat(ci): public-surface diff vs base (dropped public names, methods, test defs), advisory on PRs
The Sep 2026 whole-codebase refactor (PR #102117) opened with 1,703 public
top-level names dropped across 341 modules, 1,000 public/dunder methods in
166, and 126 `def test_` deleted in 52 files. Reviewers found ~30 of the
names by hand; the rest surfaced as post-merge rework: 10 commits restoring
symbols and facade re-exports, 6 restoring tests, and a qwen OAuth break
that passed import smoke because the caller used `module.attr`. Every one
was catchable in seconds; nothing ran the check because it did not exist.

scripts/ci/check_public_surface.py: AST diff of modules present on both
sides of merge-base..HEAD. Public top-level names (defs, classes,
assignments, imported/re-exported names), public and dunder methods of
top-level classes, and `def test_` counts per tests/ file. Deleted modules
and deleted test files are visible decisions and are not flagged; private
names are not flagged. Advisory (exit 0, prints the report) by default;
--strict exits 1 so a refactor brief or a CI lane can gate on it. Wired
into lint.yml as an advisory PR step next to the compat-pointer check.

Replayed on the refactor PR at open (63279301bcb..022785a541) it reports
exactly the figures above in 18 s; on this branch vs main it reports 0.

Test: a throwaway git repo with drops, private drops, a move-with-re-export,
a lost test def and a changed non-source module; asserts the exact report
and the advisory/strict exit codes.
2026-09-06 13:27:48 -07:00
ethernet
94779d502b Merge upstream/main into ethie/desktop-update-tests
Preserve upstream's CLI extraction and carry the installer stderr drain fix into main_install_repair alongside managed_uv.
2026-09-05 17:17:03 -04:00
yoniebans
05a18c1333 Merge upstream main (2afd17a9d7): post-simplify-codebase refactor
Two conflicts, both on the stderr-merge fix: managed_uv.py keeps the
fix on upstream's compact formatting; main.py taken from upstream (the
refactor moved _run_install_with_heartbeat to main_install_repair.py)
and the fix ported to the moved helper, which had reverted to a bare
subprocess.run.
2026-09-05 14:57:31 +02:00
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
a0be177aac fix(compat): pointers resolve to the object that MOVED, not a same-named stranger; stdin checker binds stdin= to the splatted definition
Review findings on #102117 (independent reviewer + itsflownium):

* hermes_cli.kanban_db.connect / connect_closing pointed at hermes_cli.projects_db (different DB, no
  board= parameter). The compat generator ranked candidate homes by path proximity when a name is
  defined in several modules. Now it requires shape compatibility with the BASE definition (same
  literal for constants, superset of parameter names for defs) and prefers the facade's own
  <stem>_* sibling. Same class fixed for tools.tts_tool.DEFAULT_XAI_BASE_URL (-> tts_tool_providers),
  and 17 constants/defs that had been pointed at same-named strangers (Matrix MAX_MESSAGE_LENGTH ->
  Signal's 8000, tts MAX_TEXT_LENGTH -> BlueBubbles', honcho/retaindb/supermemory *_SCHEMA -> another
  plugin's schema, ...) are now restored from BASE verbatim instead.
* send_yuanbao_direct (restored-def): body called adapter._outbound.send_direct, which HEAD moved to
  the sender; rewritten to adapter._outbound.sender.send_direct.
* COMPAT_MANIFEST.md states the scope explicitly: public top-level names only; private names and
  test monkeypatch seams are not preserved.
* scripts/check_subprocess_stdin.py: _splat_carries_stdin looked 30 lines ahead in the file text
  and was satisfied by an unrelated later stdin=; it now finds the splatted name's definition via AST
  and requires stdin inside that expression/body.

Tests: tests/test_compat_manifest_targets.py (pointer identity vs the facade's sibling; kanban
connect(board=) opens a Kanban DB, not projects.db; both FAIL on the previous layer),
test_subprocess_stdin_guard gains the false-negative probe, and the MoA -Q quiet-output contract
tests are back (tests/agent/test_moa_quiet_reference_output.py) against build_moa_facade.
2026-09-03 22:00:01 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
yoniebans
49a5c440c5 fix(install-e2e): review follow-ups — harness needle, per-matrix cap wording, typo
The harness asserted the driver still throws 'not implemented yet' for
the desktop-installer@latest update route; that arm is implemented now
(Invoke-PhaseInstallGui -Mode "update"), so the check failed against
its own tree. It asserts the implemented contract instead.

The 256-job cap wording in the workflow comment and README now states
the scope GitHub applies it at: each per-OS matrix separately, not the
combined leg count. At the 10-tag bound the largest matrix is windows
at 180.

e2e-screen-record comment: hhttps -> https.
2026-09-03 10:15:53 +02:00
yoniebans
cd39b3535c fix(install-e2e): round-2 review — decimal-only tag-count, honest cost math
^(10|[1-9])$ replaces the two-step guard: the [0-9]+ regex accepted
leading zeros that bash arithmetic then read as octal (010 passed as 8,
08 errored). README cost figures corrected to the generator's real
expansion: 41 legs/tag, 82 at the default 2 tags, update route 8/tag,
first matrix overflow at 15 tags (270 windows entries).
2026-09-02 18:32:47 +02:00
yoniebans
a3e7d6a1c7 fix(install-e2e): review findings — input hygiene, chart ranking, cost docs
tag-count now reaches the shell via the environment, validated to 1-10
(an apostrophe in the raw interpolation could terminate quoting; above
~14 tags the expansion exceeds GitHub's 256-job matrix limit).

Result-chart cell ranking matches on the leading token: rendered
success/failure cells carry artifact links, so whole-cell indexOf
ranked them -1 and any skip in the map beat a real outcome.

README documents per-run cost, route slice sizes, the tag-count bound,
and a warning against running the GUI drivers outside a disposable VM.
2026-09-02 18:27:06 +02:00
yoniebans
558d76c401 Merge upstream main (afc3d9d34c): refresh before review
One conflict: upstream 6e7c7c7da9 replaced bot-mode-closed-chat-stays-closed.spec.ts with bot-mode-row-click-mirrors-registry.spec.ts while our side had rewired its mock-server import. Kept upstream's replacement and rewired the three new specs importing ./mock-server to the consolidated tests-js copy (symbols verified present).
2026-09-02 17:53:49 +02:00
Teknium
bfbb34bbec fix(update): Windows progress server hands out its URL only once it is serving
`Start-UiServer` printed the -SelfTestUi URL (and opened the browser window)
as soon as the TcpListener was bound, but the runspace that answers /progress
starts asynchronously — BeginInvoke returns before the pipeline is open and
the script block is JIT'd, which is seconds on a loaded runner. The kernel
accepted connections into the backlog during that gap and nobody answered
them. The self-test hit it three times (#90371 and two follow-ups each
widened a timeout instead of removing the race) and it just failed an
unrelated hermes_state.py PR (run 33591547099, two 5s stale-backlog
timeouts = red).

- windows.ps1: readiness handshake after BeginInvoke — one /progress
  round-trip must succeed (≤15s) before the server is returned; on failure
  tear the listener down and continue without UI. The URL now means
  "serving", not "bound". Also fixes the browser opening to a page that never
  loads on a slow machine.
- test: 1s per-attempt probe timeout so a single dead backlog socket cannot
  consume half the readiness budget.
- CI: new `desktop_updater` classifier lane. tests/test_desktop_update_windows_*.py
  spawn the real PowerShell script; the Windows-only job now runs them only
  when scripts/desktop-update/**, the Electron updater launcher, conftest,
  pyproject, or those tests change (push/dispatch fail open). A PR that
  never touched that surface cannot be failed by its process timing.
2026-09-02 00:44:19 -07:00
Teknium
ff7745fb0a ci: every ci.yaml run 0-jobbed since 24f5a60ed1 — e2e-desktop disable as bare if: false
The `${{ false && (...) }}` if-expression on the reusable-workflow
e2e-desktop job made GitHub's workflow parser fail at startup (annotation:
"An unexpected error has occurred"), so every ci.yaml run on main and
every PR dispatched 0 jobs. Discriminator on this branch: pre-24f5a60
blob → 26 jobs; paren-relocated && form → 0; bare if: false → 25.
Lane stays disabled; comment documents the re-enable line.
2026-09-01 21:18:24 -07:00
ethernet
24f5a60ed1 fix: re-disable e2e 2026-09-01 18:08:14 -04:00
ethernet
375ce8eee5 ci: block tracked paths that collide case-insensitively
Linux is case-sensitive; Windows and macOS are not. Two tracked paths
differing only by case (README.md vs readme.md, src/Foo.py vs SRC/foo.py)
land fine on Linux and silently break every clone on a case-insensitive
host — the filesystem holds one, so checkout fails or whichever wins
clobbers the other. Git won't stop the pair from landing; it only warns
at checkout time on a case-insensitive FS. This is the enforcement point.

Adds scripts/check-case-collisions.py (index scan keyed on casefolded
full paths) + an unconditional workflow_call job wired into ci.yaml and
the all-checks-pass gate — unconditional because a collision can ship in
any kind of PR (docs, JS, config), not just Python, so gating on a
language lane would be the same passive-rule trap the infographic check
closes. Tests in tests/scripts/test_case_collision_check.py build
collisions via git update-index --cacheinfo so they run on
case-insensitive filesystems too.
2026-09-01 15:37:57 -04:00
teknium1
9ce95929a2 ci: re-enable the Desktop E2E lane — harness root-fixed by #99671
The lane was disabled Aug 2 2026 (#76627) because the mock-backend
Electron window never got a title after the Aug 1 engines/npm churn
(#76499/#76562/#76575), failing every PR identically. #99671 fixed the
root cause: per-platform/layout Electron binary resolution in the e2e
harness (apps/desktop/e2e/electron-binary.ts). The suite is green again
on Node 26 + npm 12 — delete the temporary `false &&` guard and update
the stale comment block.

Fixes #76627
2026-09-01 12:04:29 -07:00
yoniebans
d19038e76b Merge upstream main (b81383ec21) into the install-e2e suite branch
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
  the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
  one-line compat forwarder) and the new implementation already drains
  both pipes asynchronously with bounded abandonment, which supersedes
  this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
  import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
  addition inside the consolidated file; rewired the five upstream specs
  still importing './mock-server' to the consolidated path (export sets
  verified identical) and dropped the superseded apps/desktop/e2e copy.
2026-09-01 19:33:13 +02:00
yoniebans
2f25c07a2c perf(install-e2e): 60-minute job caps; 35-minute updater wait
The slowest green leg ever recorded is 29 minutes; every cap hit in the
suite's history was a hang, never work. Caps were linux 75 / macos 120 /
windows 240, so a wedged leg burned up to 4 hours of runner time to
report what its log showed in the first minutes. 60 minutes covers the
slowest leg plus cold-cache variance, and every driver-internal bound
(dmg install 45m, AHK 50m, updater wait) still fires before the job cap
in any single-hang scenario, keeping failure diagnostics specific.

The detached-updater wait drops 90m -> 35m on the same evidence: a
working updater finishes far inside 35m; a wedged one never finishes at
any bound, and the longer wait only delayed the report by an hour.
2026-09-01 17:34:13 +02:00
yoniebans
bf75c52ba2 docs(install-e2e): retire stale TODO prose now every driver arm exists
The workflow input descriptions and the skips README still declared
open-app-update and the Setup.exe re-run as driver TODOs; both run now.
Skips have exactly two causes and the prose names them: no OS entry
point for the pair, or the starting release predates the surface. The
chart's TODO label itself stays until the n/a relabel lands with the
known-broken-OLD gate work.
2026-09-01 17:11:42 +02:00
yoniebans
b7043f7669 feat(install-e2e): cover Setup.exe re-run as a windows update method
The last declared TODO: a user whose install is stale re-downloads
Hermes-Setup.exe and clicks Install over the existing install, the GUI
twin of re-running the one-liner. Windows shows the full installer UI on
a re-run (the already-installed fast path is macOS-only), so the existing
AHK install drive applies unchanged; install.ps1's repository stage
fetches the existing checkout forward to what main serves, now HEAD.

Invoke-PhaseInstallGui gains an update mode instead of a parallel copy:
the phase label, proof dir, and expected-sha assertion become parameters,
and the update-is-available assert stays install-only. The bootstrap log
rotates before the re-run so the AHK's completion fallback cannot match
the install phase's old completion line.
2026-09-01 15:03:59 +02:00