`generate-skill-docs.py` now deletes every page under `user-guide/skills/{bundled,optional}`
it did not write this run, together with the zh-Hans mirror twin, so a skill that moves,
merges or leaves the shipped set takes its page with it instead of lingering as an orphan
that cross-links still reach (24 such pages after the shipped-set slim, plus 21 zh-Hans
copies whose English page was already gone).
The Docs Site Checks workflow regenerated the docs and never compared the result with the
committed copies, which is what GitHub renders; it now fails with a pointer to the generator
when they differ. Also regenerates the two pages that drifted since the salvaged commits.
scripts/check_no_tmp_literals.py flags /tmp path tokens in production code, skills,
docs and prompt strings (tests, CI workflows, Dockerfiles, lockfiles, i18n mirror,
code comments and docstrings exempt; ${TMPDIR:-/tmp} idiom exempt). Opt out one line
with 'no-tmp: ok — <why>' on the line or the line above. _BASELINE lists pre-existing
hits per file: growth fails, burn-down is advisory (--strict-baseline / --print-baseline
to refresh). Wired into lint.yml next to check_compat_pointers.
Install-Uv accepted any file at $HermesHome\bin\uv.exe, and copied whatever
`Get-Command uv` returned into that location. Chocolatey's bin\uv.exe is a
ShimGen launcher that locates ..\lib\uv\tools\uv.exe RELATIVE to itself, so
the copy is dead on arrival; `& exe --version` does not throw on a nonzero
exit, so the launcher passed the try/catch and the Python stage then failed
with "Python 3.11 not available" (#110350). The re-run path trusted the same
broken copy again.
Building on KoNit-K's Test-ManagedUvBinary and its three call sites:
- Test-ManagedUvBinary merges stderr, relaxes the error preference, and
returns the `uv <version>` line only on exit 0 -- a launcher's error text
can no longer surface as "Managed uv found (Cannot find file ...)".
- Resolve-UvShimTarget maps a candidate to the standalone binary before the
copy: `<name>.shim` sidecar (Scoop), the Chocolatey bin\ -> lib\<pkg>\tools\
layout, symlinks (winget Links\); other reparse points (WindowsApps
app-execution aliases) have no copyable file and skip the salvage.
- The salvage rung validates the candidate where it lives, copies, then
validates the COPY at its new location and removes it on failure, so the
stage fails honestly instead of reporting success over a dead launcher.
- scripts/tests/test-install-ps1-uv-shim-validation.ps1 drives the real
Install-Uv with compiled fake uv binaries (a working uv and a
location-relative launcher) under stubbed installer rungs; wired into
installer-tests.yml for pwsh 7 and Windows PowerShell 5.1.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
`run-workspace-checks.mjs` buffers each check's output, writes it in one
piece, and then called `process.exit(1)`. On a pipe, stdout is asynchronous,
so the exit dropped what was still buffered: the log of a failing check
stopped in the middle, before the test runner's failure summary.
Set `process.exitCode` and return, so the process ends after stdout drains.
Docs under website/docs were linked with Docusaurus site routes
(`](/getting-started/installation)`, `](/docs/user-guide/x)`), which GitHub's
file viewer resolves as repository paths and 404s (#114428). Relative Markdown
file links (`](../user-guide/x.md#anchor)`) are followed by both GitHub and
Docusaurus, so that becomes the authoring convention:
- website/scripts/check_doc_links.py lints hand-authored EN + zh-Hans pages
for route-style links (`--fix` rewrites them, refusing any route that maps
to no doc file); wired into the Docs Site Checks workflow and
tests/website/test_check_doc_links.py.
- generate-skill-docs.py emits the same relative form for related-skill and
catalog links instead of `/docs/user-guide/skills/...`.
- src/remark/relativeDocLinks.js rewrites `./x.md`/`../x.md` to
content-root-absolute `/x.md` before Docusaurus resolves links, so a
relative link still resolves in the zh-Hans build when source and target
sit on different sides of the translation fallback (Docusaurus resolves
`./`/`../` only against the source file's own directory).
- website/README.md states the convention and points at the checker.
Trim the salvaged retry to the case the logs show: `_get_json` returning None
(timeout / non-200) is a hole in the cursor walk, not the end of the catalog.
A dict page with an empty item list stays terminal as before, and the
interactive browse path (12 s budget) no longer sleeps up to 8 s per retry.
Constant `CATALOG_PAGE_RETRIES` replaces the local literal.
Workflow: `skills-index.yml` `timeout-minutes` 50 -> 120. Since Sep 17 ClawHub
serves ~10 s per 200-item page (was ~2.5 s), so the ~400-page walk alone takes
60-70 min; the 06:28 UTC runs on Sep 17 and Sep 18 were cancelled at 50 min with
the clawhub crawl still running, and the 18:19 run stopped at 12,145 skills
(< 20,000 floor) after one failed page. Both halves are needed for the index to
ship again.
Catalog CI validates each pinned tree in a venv that has only hermes-agent, so any plugin whose
code arrives through pyproject dependencies (the wrapper shape from #113851) failed the probe with
"No module named ...". The flag runs the same constrained install `plugins install` would, then
probes; the workflow passes it. Live: mnemosyne pin fails without the flag, passes with it.
The self-lock/holder preflight (#99711) deferred the repair on the theory
that the updater's own mapped python.exe makes the venv rename impossible.
Live on windows-latest, a process executing from venv\Scripts\python.exe
(and the repo's real .venv with cp311 .pyd extensions loaded) does NOT block
the rename; what blocks it with WinError 5 is any ordinary handle under the
tree: a process cwd, an open file, a sync client. Since sys.executable is
always under the venv on Windows, the deferral fired on every run and no
Windows install could repair from `hermes update`. Retire it and its tests;
the pyvenv.cfg repoint needs no rename and has none of that exposure. The
wine2e lane now runs the cutover probes instead.
The repoint keeps the live venv's site-packages, so provisioning's
fall-forward to the next minor (3.11 -> 3.12, #76106) must not be pointed at
a cp311 tree: refuse before touching pyvenv.cfg. The live Windows test holds
the venv the way the field does (a child with cwd inside it), asserts the
rename path fails (the symptom) and the repoint succeeds; the rollback and
minor-guard tests run on every host.
scripts/ci/check_os_marker_fakes.py flags test files that make the
interpreter believe it is on macOS (is_macos -> True, sys.platform ->
"darwin", platform.system -> "Darwin") while carrying no `macos_only`
marker: the macOS lane imports only marked files, so such a file is green
on Linux over a faked branch and never runs on the host it exists for.
A `# os-marker: ok — <why>` comment opts a host-independent line out.
_BASELINE holds the files that already faked macOS when the check landed;
a stale entry fails the check so the list can only burn down. Wired into
lint.yml next to the compat-pointer check; two invariant tests cover a
flagged fake vs marked/opted-out files and host-honest platform reads.
The zsh login-shell legs in remote-lifecycle.test.ts and
ssh-connection.test.ts silently returned when zsh was missing, and the
js-tests runner image ships no zsh, so the #111949 coverage never ran on
CI and a wrapper regression stayed green.
- js-tests.yml: install zsh on the Linux runner before the checks.
- Both legs now report vitest skips ('zsh not installed') instead of
passing; the ssh-connection leg is its own test so the skip is visible.
- Docs: note the zsh degraded mode (no process-group kill for a hung
probe's grandchildren) in the SSH connection guide.
The scan is detection-only and its findings are the repo-wide baseline of
CVEs in pinned dependencies — identical for every PR, unrelated to any
PR's diff. Reporting that baseline in each PR's review comment read as
"this PR has 76 vulnerabilities" to contributors, and the SARIF upload
tripped GitHub's per-installation API rate limit during merge trains.
The scheduled weekly run (plus workflow_dispatch) keeps feeding the
Security tab; the per-PR workflow_call, the review_status wrapper job,
and the orchestrator's now-unneeded SARIF permissions are removed.
scripts/check_profile_scope_patterns.py runs the validated hazard regexes in
scripts/ci/profile_scope_patterns.json (18 of the 31 campaign patterns: every one has a scope_hint
and hits <= 50 sites on main; the wider ones are review greps, not lint) against the lines added
vs the PR base and prints file:line, pattern id/class and why. Always exits 0: most shapes have
legitimate sites (a standalone `hermes -p x` process where environ IS the profile), so the
reviewer reads each finding against its scope hint. Wired into lint.yml beside the public-surface
diff with continue-on-error.
Proof: the pre-fix tools/bot_relay.py (`env = dict(os.environ)`, before the served_profile_child_env
change) is flagged as P05/C2; the fixed file and this branch's diff vs main report 0 findings.
Test: a fixture with the hazard is flagged on the right lines, the scoped-builder version is not,
and the line filter hides hits outside the added range.
Aligns the star ranking with the rule the skills index already follows:
GitHub is consulted only by the twice-daily skills-index.yml schedule, whose
artifact every docs deploy reuses. deploy-site.yml now runs
fetch-plugin-stars.py without --probe (reuse-only: artifact → live site copy →
disk → empty) and cannot call the API at all, so a same-day merge train adds
zero requests regardless of cache age.
The probe itself collapses from one REST call per repo (~26 today, growing
with the catalog) to a single GraphQL query with aliased repository fields,
so the scheduled run costs one request no matter how big the catalog gets. A
failed probe (rate limit, renamed repo, bad token) keeps the previous counts.
Greptile's two findings on the original PR were both right.
1. The scaler read test_durations.json from the checkout, but CI ran on
a fresh runner where that file never exists (it is gitignored and the
slicing-era artifact/merge job that produced it is gone). The feature
was inert exactly where the false FLAKY kills happen. tests.yml now
restores the most recent main-saved cache before the run (PRs read
only) and saves it after a green push to main, mirroring the
ci-timings-baseline restore/save pattern already in ci.yaml.
2. _save_durations persisted every file's total subprocess wall,
including the ~cap of a timed-out attempt and the retry-summed wall
of a FLAKY file. With the scaler that compounds: a hang cached at
~300s earns 900s next run, then ~900s cached earns 2700s, until the
job timeout is the only bound. _clean_pass_durations drops failed and
FLAKY files from the write so a file's cached duration is always a
first-attempt-clean measurement; those files keep their previous
known-good entry.
Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
Catalog entries sort official → stars desc → name, both in browse shelves and
filtered grids, with a ★ pill on each card linking to the repo's stargazers.
Rate-limit discipline is the design constraint: the docs site deploys many
times a day and shares one GitHub App API budget with every other workflow
(tonight's merge train got rate-limited on unrelated uploads). So
website/scripts/fetch-plugin-stars.py first fetches the live site's own
plugin-stars.json (a CDN GET, not the API); if that cache is under 24h old it
is reused verbatim and GitHub is never called. Only a stale cache triggers one
GET /repos/{owner}/{repo} per unique catalog repo, and a 403/429 mid-run keeps
the previous counts instead of zeroing them. extract-plugins.py merges the
cache into plugins.json (`stars`) and plugins-meta.json (`starsFetchedAt`), and
the page footnote says when the ranking was last refreshed.
osv-scanner.yml documents itself as detection-only (fail-on-vuln: false,
findings land in the Security tab) yet all-checks-pass listed it in needs,
so any failure result blocked the merge. In practice the failures are not
vulnerabilities: the "Upload to code-scanning" step hits GitHub's
per-installation API rate limit whenever several PRs run at once, and a
merge train of catalog entries went red on it across the board. The scan
still runs on every PR and weekly on main; it just reports instead of gating.
Teknium's ruling: catalog plugins do not need the 2-week maturity window,
but they may not ship an in-app updater that downloads and replaces their
own files, because that makes the reviewed SHA pin decorative. Rule 3 in
the README and item 5 on the docs page now say so, and plugin-catalog-ci
fails an entry whose catalog build both fetches from GitHub releases/raw
and writes or renames plugin files (either half alone is allowed).
IronClaw's #7756 swept every unbounded CI operation (apt hangs, uncapped
jobs, external downloads). Same sweep here found exactly one gap: the
osv-scanner emit-status wrapper job had no timeout-minutes, so a wedged
artifact download could hold a runner for GitHub's 6-hour default. Every
other job across all 30 workflows is already bounded. Capped at 10m.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute
job timeout, so the live skills-index.json has been frozen at that date and every
`hermes skills search` fell through to live GitHub API calls (~500 inspect requests per
cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been
appending to #66616 four times a day since.
Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to
fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k
skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it.
- enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder
without an owner (the link is a nicety; the index is not).
- build_skills_index.py passes an 8-minute budget.
- skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path
(clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min).
The advisory step ran `git merge-base origin/main HEAD` where HEAD is the
checked-out refs/pull/N/merge commit. GitHub recomputes that synthetic merge
commit whenever main moves, so by the time the job runs the checked-out SHA is
often no longer reachable from any ref. `git fetch --deepen` on an
unreachable commit can never pull in its parents, so all four fetches in the
loop (200 + 3 x 1000) ran to completion, burned 4m15s, and still ended in
"public-surface: no merge-base" (run 34285107265, job 102259026802). The
step-level timeout keeps that overrun from cancelling the blocking job; this
commit makes the step finish on the first fetch instead.
Fetch refs/pull/N/head explicitly (always reachable server-side) and diff
that ref against the base; the merge commit's own content is irrelevant to a
symbol diff.
Raise the step timeout from 2 to 3 minutes: one --deepen=200 fetch took ~56s
and a --deepen=1000 ~67s on the runner, and main gains ~170 commits a day, so a
branch a day old legitimately needs both. 45s of blocking steps + 3 min still
fits the 5-minute job budget.
The 'Public-surface diff vs base (advisory)' step already carries
continue-on-error: true, so its own failure never fails the Windows-footguns
job. But continue-on-error does not shield the job-level timeout-minutes: 5:
when the deepen/fetch loop runs long (a distant or missing merge-base), the
step eats the whole job budget and the job is cancelled at ~5m even though the
blocking footgun/compat checks passed and the public-surface check is advisory.
Give the advisory step its own timeout-minutes: 2. A step overrun is then
killed and, via continue-on-error, kept off the job outcome the same way a
step failure already is, so the blocking job can finish under its budget.
Fixes#106103.
- hermes_cli/plugins_cmd_catalog.py: new sibling owning resolution, the
.hermes-catalog.json provenance sidecar, search/info/validate, re-pin on
update, and the dashboard/TUI payload builders. plugins_cmd.py only
gains the hooks (cmd_install catalog branch, cmd_update / dashboard
update re-pin, dashboard_install_plugin catalog_name + kill list,
dispatch entries); the community index (plugin_index.py) is gone.
- hermes_cli/plugin_catalog.py: catalog_dir parameter replaces the
test-only HERMES_PLUGIN_CATALOG_DIR env var; live refresh reads ONE
published document (/docs/api/plugin-catalog.json, 6h cache, in-tree
fallback) instead of the unauthenticated GitHub contents API (60 req/h,
1 request per entry); in-tree and live removals are unioned so a stale
cache can never un-block.
- Catalog route lives in web_routers/dashboard_ui.py (the facade is off
limits); _plugin_runtime_status shared from web_server_dashboard.py;
hub rows carry removed_reason. TUI plugins.manage gains catalog_name
install, catalog row fields and an update action.
- plugin_validate: the probe context honours ctx.get_config defaults
(real plugins do int(ctx.get_config("timeout", 180)) in register()).
- Installed-state merge matches through the sidecar's catalog_name
first — catalog names rarely equal manifest names.
- extract-plugins.py emits plugin-catalog.json; deploy-site triggers on
plugin-catalog/** so entry merges republish it.
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
Both sides fixed the swallowed-click class on the onboarding picker.
Hers is the root cause: persistent 100% zoom through the app's own
setting, verified from both the renderer IPC and the BrowserWindow, and
re-applied before the dismiss loop — scale drift is what moved real
click points onto the wrapping container. The dispatchEvent fallback is
dropped: it bypassed hit-testing, so a leg could pass where a real
user's click would fail.
Comment-only conflicts in managed_uv.py and main_install_repair.py
resolved by keeping the fuller mechanism text (import-order reach and
the legacy hand-off scope).
Keep unknown failures red, rotate evidence per attempt, and emit receipts for signature-confirmed historical cases. Add CI-only diagnostics and an exact-tag input for the unresolved July hand-off.
The round-2 change made check_public_surface refuse to report a clean diff
without a merge-base (exit 2). That correctly exposed that the lint job's
depth-1 checkout plus a depth-1 fetch of the base has NO merge-base, so the
advisory step had been silently reporting 0 drops on every PR. The step now
deepens both sides until a merge-base exists and carries continue-on-error so
an advisory check can never block the Windows-footguns job it rides in.
The Sep 2026 whole-codebase refactor (PR #102117) opened with 1,703 public
top-level names dropped across 341 modules, 1,000 public/dunder methods in
166, and 126 `def test_` deleted in 52 files. Reviewers found ~30 of the
names by hand; the rest surfaced as post-merge rework: 10 commits restoring
symbols and facade re-exports, 6 restoring tests, and a qwen OAuth break
that passed import smoke because the caller used `module.attr`. Every one
was catchable in seconds; nothing ran the check because it did not exist.
scripts/ci/check_public_surface.py: AST diff of modules present on both
sides of merge-base..HEAD. Public top-level names (defs, classes,
assignments, imported/re-exported names), public and dunder methods of
top-level classes, and `def test_` counts per tests/ file. Deleted modules
and deleted test files are visible decisions and are not flagged; private
names are not flagged. Advisory (exit 0, prints the report) by default;
--strict exits 1 so a refactor brief or a CI lane can gate on it. Wired
into lint.yml as an advisory PR step next to the compat-pointer check.
Replayed on the refactor PR at open (63279301bcb..022785a541) it reports
exactly the figures above in 18 s; on this branch vs main it reports 0.
Test: a throwaway git repo with drops, private drops, a move-with-re-export,
a lost test def and a changed non-source module; asserts the exact report
and the advisory/strict exit codes.
Two conflicts, both on the stderr-merge fix: managed_uv.py keeps the
fix on upstream's compact formatting; main.py taken from upstream (the
refactor moved _run_install_with_heartbeat to main_install_repair.py)
and the fix ported to the moved helper, which had reverted to a bare
subprocess.run.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
Review findings on #102117 (independent reviewer + itsflownium):
* hermes_cli.kanban_db.connect / connect_closing pointed at hermes_cli.projects_db (different DB, no
board= parameter). The compat generator ranked candidate homes by path proximity when a name is
defined in several modules. Now it requires shape compatibility with the BASE definition (same
literal for constants, superset of parameter names for defs) and prefers the facade's own
<stem>_* sibling. Same class fixed for tools.tts_tool.DEFAULT_XAI_BASE_URL (-> tts_tool_providers),
and 17 constants/defs that had been pointed at same-named strangers (Matrix MAX_MESSAGE_LENGTH ->
Signal's 8000, tts MAX_TEXT_LENGTH -> BlueBubbles', honcho/retaindb/supermemory *_SCHEMA -> another
plugin's schema, ...) are now restored from BASE verbatim instead.
* send_yuanbao_direct (restored-def): body called adapter._outbound.send_direct, which HEAD moved to
the sender; rewritten to adapter._outbound.sender.send_direct.
* COMPAT_MANIFEST.md states the scope explicitly: public top-level names only; private names and
test monkeypatch seams are not preserved.
* scripts/check_subprocess_stdin.py: _splat_carries_stdin looked 30 lines ahead in the file text
and was satisfied by an unrelated later stdin=; it now finds the splatted name's definition via AST
and requires stdin inside that expression/body.
Tests: tests/test_compat_manifest_targets.py (pointer identity vs the facade's sibling; kanban
connect(board=) opens a Kanban DB, not projects.db; both FAIL on the previous layer),
test_subprocess_stdin_guard gains the false-negative probe, and the MoA -Q quiet-output contract
tests are back (tests/agent/test_moa_quiet_reference_output.py) against build_moa_facade.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
The harness asserted the driver still throws 'not implemented yet' for
the desktop-installer@latest update route; that arm is implemented now
(Invoke-PhaseInstallGui -Mode "update"), so the check failed against
its own tree. It asserts the implemented contract instead.
The 256-job cap wording in the workflow comment and README now states
the scope GitHub applies it at: each per-OS matrix separately, not the
combined leg count. At the 10-tag bound the largest matrix is windows
at 180.
e2e-screen-record comment: hhttps -> https.
^(10|[1-9])$ replaces the two-step guard: the [0-9]+ regex accepted
leading zeros that bash arithmetic then read as octal (010 passed as 8,
08 errored). README cost figures corrected to the generator's real
expansion: 41 legs/tag, 82 at the default 2 tags, update route 8/tag,
first matrix overflow at 15 tags (270 windows entries).
tag-count now reaches the shell via the environment, validated to 1-10
(an apostrophe in the raw interpolation could terminate quoting; above
~14 tags the expansion exceeds GitHub's 256-job matrix limit).
Result-chart cell ranking matches on the leading token: rendered
success/failure cells carry artifact links, so whole-cell indexOf
ranked them -1 and any skip in the map beat a real outcome.
README documents per-run cost, route slice sizes, the tag-count bound,
and a warning against running the GUI drivers outside a disposable VM.
One conflict: upstream 6e7c7c7da9 replaced bot-mode-closed-chat-stays-closed.spec.ts with bot-mode-row-click-mirrors-registry.spec.ts while our side had rewired its mock-server import. Kept upstream's replacement and rewired the three new specs importing ./mock-server to the consolidated tests-js copy (symbols verified present).
`Start-UiServer` printed the -SelfTestUi URL (and opened the browser window)
as soon as the TcpListener was bound, but the runspace that answers /progress
starts asynchronously — BeginInvoke returns before the pipeline is open and
the script block is JIT'd, which is seconds on a loaded runner. The kernel
accepted connections into the backlog during that gap and nobody answered
them. The self-test hit it three times (#90371 and two follow-ups each
widened a timeout instead of removing the race) and it just failed an
unrelated hermes_state.py PR (run 33591547099, two 5s stale-backlog
timeouts = red).
- windows.ps1: readiness handshake after BeginInvoke — one /progress
round-trip must succeed (≤15s) before the server is returned; on failure
tear the listener down and continue without UI. The URL now means
"serving", not "bound". Also fixes the browser opening to a page that never
loads on a slow machine.
- test: 1s per-attempt probe timeout so a single dead backlog socket cannot
consume half the readiness budget.
- CI: new `desktop_updater` classifier lane. tests/test_desktop_update_windows_*.py
spawn the real PowerShell script; the Windows-only job now runs them only
when scripts/desktop-update/**, the Electron updater launcher, conftest,
pyproject, or those tests change (push/dispatch fail open). A PR that
never touched that surface cannot be failed by its process timing.
The `${{ false && (...) }}` if-expression on the reusable-workflow
e2e-desktop job made GitHub's workflow parser fail at startup (annotation:
"An unexpected error has occurred"), so every ci.yaml run on main and
every PR dispatched 0 jobs. Discriminator on this branch: pre-24f5a60
blob → 26 jobs; paren-relocated && form → 0; bare if: false → 25.
Lane stays disabled; comment documents the re-enable line.
Linux is case-sensitive; Windows and macOS are not. Two tracked paths
differing only by case (README.md vs readme.md, src/Foo.py vs SRC/foo.py)
land fine on Linux and silently break every clone on a case-insensitive
host — the filesystem holds one, so checkout fails or whichever wins
clobbers the other. Git won't stop the pair from landing; it only warns
at checkout time on a case-insensitive FS. This is the enforcement point.
Adds scripts/check-case-collisions.py (index scan keyed on casefolded
full paths) + an unconditional workflow_call job wired into ci.yaml and
the all-checks-pass gate — unconditional because a collision can ship in
any kind of PR (docs, JS, config), not just Python, so gating on a
language lane would be the same passive-rule trap the infographic check
closes. Tests in tests/scripts/test_case_collision_check.py build
collisions via git update-index --cacheinfo so they run on
case-insensitive filesystems too.
The lane was disabled Aug 2 2026 (#76627) because the mock-backend
Electron window never got a title after the Aug 1 engines/npm churn
(#76499/#76562/#76575), failing every PR identically. #99671 fixed the
root cause: per-platform/layout Electron binary resolution in the e2e
harness (apps/desktop/e2e/electron-binary.ts). The suite is green again
on Node 26 + npm 12 — delete the temporary `false &&` guard and update
the stale comment block.
Fixes#76627
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
one-line compat forwarder) and the new implementation already drains
both pipes asynchronously with bounded abandonment, which supersedes
this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
addition inside the consolidated file; rewired the five upstream specs
still importing './mock-server' to the consolidated path (export sets
verified identical) and dropped the superseded apps/desktop/e2e copy.
The slowest green leg ever recorded is 29 minutes; every cap hit in the
suite's history was a hang, never work. Caps were linux 75 / macos 120 /
windows 240, so a wedged leg burned up to 4 hours of runner time to
report what its log showed in the first minutes. 60 minutes covers the
slowest leg plus cold-cache variance, and every driver-internal bound
(dmg install 45m, AHK 50m, updater wait) still fires before the job cap
in any single-hang scenario, keeping failure diagnostics specific.
The detached-updater wait drops 90m -> 35m on the same evidence: a
working updater finishes far inside 35m; a wedged one never finishes at
any bound, and the longer wait only delayed the report by an hour.
The workflow input descriptions and the skips README still declared
open-app-update and the Setup.exe re-run as driver TODOs; both run now.
Skips have exactly two causes and the prose names them: no OS entry
point for the pair, or the starting release predates the surface. The
chart's TODO label itself stays until the n/a relabel lands with the
known-broken-OLD gate work.
The last declared TODO: a user whose install is stale re-downloads
Hermes-Setup.exe and clicks Install over the existing install, the GUI
twin of re-running the one-liner. Windows shows the full installer UI on
a re-run (the already-installed fast path is macOS-only), so the existing
AHK install drive applies unchanged; install.ps1's repository stage
fetches the existing checkout forward to what main serves, now HEAD.
Invoke-PhaseInstallGui gains an update mode instead of a parallel copy:
the phase label, proof dir, and expected-sha assertion become parameters,
and the update-is-available assert stays install-only. The bootstrap log
rotates before the re-run so the AHK's completion fallback cannot match
the install phase's old completion line.