scripts/ci/check_os_marker_fakes.py flags test files that make the
interpreter believe it is on macOS (is_macos -> True, sys.platform ->
"darwin", platform.system -> "Darwin") while carrying no `macos_only`
marker: the macOS lane imports only marked files, so such a file is green
on Linux over a faked branch and never runs on the host it exists for.
A `# os-marker: ok — <why>` comment opts a host-independent line out.
_BASELINE holds the files that already faked macOS when the check landed;
a stale entry fails the check so the list can only burn down. Wired into
lint.yml next to the compat-pointer check; two invariant tests cover a
flagged fake vs marked/opted-out files and host-honest platform reads.
The zsh login-shell legs in remote-lifecycle.test.ts and
ssh-connection.test.ts silently returned when zsh was missing, and the
js-tests runner image ships no zsh, so the #111949 coverage never ran on
CI and a wrapper regression stayed green.
- js-tests.yml: install zsh on the Linux runner before the checks.
- Both legs now report vitest skips ('zsh not installed') instead of
passing; the ssh-connection leg is its own test so the skip is visible.
- Docs: note the zsh degraded mode (no process-group kill for a hung
probe's grandchildren) in the SSH connection guide.
The scan is detection-only and its findings are the repo-wide baseline of
CVEs in pinned dependencies — identical for every PR, unrelated to any
PR's diff. Reporting that baseline in each PR's review comment read as
"this PR has 76 vulnerabilities" to contributors, and the SARIF upload
tripped GitHub's per-installation API rate limit during merge trains.
The scheduled weekly run (plus workflow_dispatch) keeps feeding the
Security tab; the per-PR workflow_call, the review_status wrapper job,
and the orchestrator's now-unneeded SARIF permissions are removed.
scripts/check_profile_scope_patterns.py runs the validated hazard regexes in
scripts/ci/profile_scope_patterns.json (18 of the 31 campaign patterns: every one has a scope_hint
and hits <= 50 sites on main; the wider ones are review greps, not lint) against the lines added
vs the PR base and prints file:line, pattern id/class and why. Always exits 0: most shapes have
legitimate sites (a standalone `hermes -p x` process where environ IS the profile), so the
reviewer reads each finding against its scope hint. Wired into lint.yml beside the public-surface
diff with continue-on-error.
Proof: the pre-fix tools/bot_relay.py (`env = dict(os.environ)`, before the served_profile_child_env
change) is flagged as P05/C2; the fixed file and this branch's diff vs main report 0 findings.
Test: a fixture with the hazard is flagged on the right lines, the scoped-builder version is not,
and the line filter hides hits outside the added range.
Aligns the star ranking with the rule the skills index already follows:
GitHub is consulted only by the twice-daily skills-index.yml schedule, whose
artifact every docs deploy reuses. deploy-site.yml now runs
fetch-plugin-stars.py without --probe (reuse-only: artifact → live site copy →
disk → empty) and cannot call the API at all, so a same-day merge train adds
zero requests regardless of cache age.
The probe itself collapses from one REST call per repo (~26 today, growing
with the catalog) to a single GraphQL query with aliased repository fields,
so the scheduled run costs one request no matter how big the catalog gets. A
failed probe (rate limit, renamed repo, bad token) keeps the previous counts.
Greptile's two findings on the original PR were both right.
1. The scaler read test_durations.json from the checkout, but CI ran on
a fresh runner where that file never exists (it is gitignored and the
slicing-era artifact/merge job that produced it is gone). The feature
was inert exactly where the false FLAKY kills happen. tests.yml now
restores the most recent main-saved cache before the run (PRs read
only) and saves it after a green push to main, mirroring the
ci-timings-baseline restore/save pattern already in ci.yaml.
2. _save_durations persisted every file's total subprocess wall,
including the ~cap of a timed-out attempt and the retry-summed wall
of a FLAKY file. With the scaler that compounds: a hang cached at
~300s earns 900s next run, then ~900s cached earns 2700s, until the
job timeout is the only bound. _clean_pass_durations drops failed and
FLAKY files from the write so a file's cached duration is always a
first-attempt-clean measurement; those files keep their previous
known-good entry.
Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
Catalog entries sort official → stars desc → name, both in browse shelves and
filtered grids, with a ★ pill on each card linking to the repo's stargazers.
Rate-limit discipline is the design constraint: the docs site deploys many
times a day and shares one GitHub App API budget with every other workflow
(tonight's merge train got rate-limited on unrelated uploads). So
website/scripts/fetch-plugin-stars.py first fetches the live site's own
plugin-stars.json (a CDN GET, not the API); if that cache is under 24h old it
is reused verbatim and GitHub is never called. Only a stale cache triggers one
GET /repos/{owner}/{repo} per unique catalog repo, and a 403/429 mid-run keeps
the previous counts instead of zeroing them. extract-plugins.py merges the
cache into plugins.json (`stars`) and plugins-meta.json (`starsFetchedAt`), and
the page footnote says when the ranking was last refreshed.
osv-scanner.yml documents itself as detection-only (fail-on-vuln: false,
findings land in the Security tab) yet all-checks-pass listed it in needs,
so any failure result blocked the merge. In practice the failures are not
vulnerabilities: the "Upload to code-scanning" step hits GitHub's
per-installation API rate limit whenever several PRs run at once, and a
merge train of catalog entries went red on it across the board. The scan
still runs on every PR and weekly on main; it just reports instead of gating.
Teknium's ruling: catalog plugins do not need the 2-week maturity window,
but they may not ship an in-app updater that downloads and replaces their
own files, because that makes the reviewed SHA pin decorative. Rule 3 in
the README and item 5 on the docs page now say so, and plugin-catalog-ci
fails an entry whose catalog build both fetches from GitHub releases/raw
and writes or renames plugin files (either half alone is allowed).
IronClaw's #7756 swept every unbounded CI operation (apt hangs, uncapped
jobs, external downloads). Same sweep here found exactly one gap: the
osv-scanner emit-status wrapper job had no timeout-minutes, so a wedged
artifact download could hold a runner for GitHub's 6-hour default. Every
other job across all 30 workflows is already bounded. Capped at 10m.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute
job timeout, so the live skills-index.json has been frozen at that date and every
`hermes skills search` fell through to live GitHub API calls (~500 inspect requests per
cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been
appending to #66616 four times a day since.
Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to
fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k
skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it.
- enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder
without an owner (the link is a nicety; the index is not).
- build_skills_index.py passes an 8-minute budget.
- skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path
(clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min).
The advisory step ran `git merge-base origin/main HEAD` where HEAD is the
checked-out refs/pull/N/merge commit. GitHub recomputes that synthetic merge
commit whenever main moves, so by the time the job runs the checked-out SHA is
often no longer reachable from any ref. `git fetch --deepen` on an
unreachable commit can never pull in its parents, so all four fetches in the
loop (200 + 3 x 1000) ran to completion, burned 4m15s, and still ended in
"public-surface: no merge-base" (run 34285107265, job 102259026802). The
step-level timeout keeps that overrun from cancelling the blocking job; this
commit makes the step finish on the first fetch instead.
Fetch refs/pull/N/head explicitly (always reachable server-side) and diff
that ref against the base; the merge commit's own content is irrelevant to a
symbol diff.
Raise the step timeout from 2 to 3 minutes: one --deepen=200 fetch took ~56s
and a --deepen=1000 ~67s on the runner, and main gains ~170 commits a day, so a
branch a day old legitimately needs both. 45s of blocking steps + 3 min still
fits the 5-minute job budget.
The 'Public-surface diff vs base (advisory)' step already carries
continue-on-error: true, so its own failure never fails the Windows-footguns
job. But continue-on-error does not shield the job-level timeout-minutes: 5:
when the deepen/fetch loop runs long (a distant or missing merge-base), the
step eats the whole job budget and the job is cancelled at ~5m even though the
blocking footgun/compat checks passed and the public-surface check is advisory.
Give the advisory step its own timeout-minutes: 2. A step overrun is then
killed and, via continue-on-error, kept off the job outcome the same way a
step failure already is, so the blocking job can finish under its budget.
Fixes#106103.
- hermes_cli/plugins_cmd_catalog.py: new sibling owning resolution, the
.hermes-catalog.json provenance sidecar, search/info/validate, re-pin on
update, and the dashboard/TUI payload builders. plugins_cmd.py only
gains the hooks (cmd_install catalog branch, cmd_update / dashboard
update re-pin, dashboard_install_plugin catalog_name + kill list,
dispatch entries); the community index (plugin_index.py) is gone.
- hermes_cli/plugin_catalog.py: catalog_dir parameter replaces the
test-only HERMES_PLUGIN_CATALOG_DIR env var; live refresh reads ONE
published document (/docs/api/plugin-catalog.json, 6h cache, in-tree
fallback) instead of the unauthenticated GitHub contents API (60 req/h,
1 request per entry); in-tree and live removals are unioned so a stale
cache can never un-block.
- Catalog route lives in web_routers/dashboard_ui.py (the facade is off
limits); _plugin_runtime_status shared from web_server_dashboard.py;
hub rows carry removed_reason. TUI plugins.manage gains catalog_name
install, catalog row fields and an update action.
- plugin_validate: the probe context honours ctx.get_config defaults
(real plugins do int(ctx.get_config("timeout", 180)) in register()).
- Installed-state merge matches through the sidecar's catalog_name
first — catalog names rarely equal manifest names.
- extract-plugins.py emits plugin-catalog.json; deploy-site triggers on
plugin-catalog/** so entry merges republish it.
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
Both sides fixed the swallowed-click class on the onboarding picker.
Hers is the root cause: persistent 100% zoom through the app's own
setting, verified from both the renderer IPC and the BrowserWindow, and
re-applied before the dismiss loop — scale drift is what moved real
click points onto the wrapping container. The dispatchEvent fallback is
dropped: it bypassed hit-testing, so a leg could pass where a real
user's click would fail.
Comment-only conflicts in managed_uv.py and main_install_repair.py
resolved by keeping the fuller mechanism text (import-order reach and
the legacy hand-off scope).
Keep unknown failures red, rotate evidence per attempt, and emit receipts for signature-confirmed historical cases. Add CI-only diagnostics and an exact-tag input for the unresolved July hand-off.
The round-2 change made check_public_surface refuse to report a clean diff
without a merge-base (exit 2). That correctly exposed that the lint job's
depth-1 checkout plus a depth-1 fetch of the base has NO merge-base, so the
advisory step had been silently reporting 0 drops on every PR. The step now
deepens both sides until a merge-base exists and carries continue-on-error so
an advisory check can never block the Windows-footguns job it rides in.
The Sep 2026 whole-codebase refactor (PR #102117) opened with 1,703 public
top-level names dropped across 341 modules, 1,000 public/dunder methods in
166, and 126 `def test_` deleted in 52 files. Reviewers found ~30 of the
names by hand; the rest surfaced as post-merge rework: 10 commits restoring
symbols and facade re-exports, 6 restoring tests, and a qwen OAuth break
that passed import smoke because the caller used `module.attr`. Every one
was catchable in seconds; nothing ran the check because it did not exist.
scripts/ci/check_public_surface.py: AST diff of modules present on both
sides of merge-base..HEAD. Public top-level names (defs, classes,
assignments, imported/re-exported names), public and dunder methods of
top-level classes, and `def test_` counts per tests/ file. Deleted modules
and deleted test files are visible decisions and are not flagged; private
names are not flagged. Advisory (exit 0, prints the report) by default;
--strict exits 1 so a refactor brief or a CI lane can gate on it. Wired
into lint.yml as an advisory PR step next to the compat-pointer check.
Replayed on the refactor PR at open (63279301bcb..022785a541) it reports
exactly the figures above in 18 s; on this branch vs main it reports 0.
Test: a throwaway git repo with drops, private drops, a move-with-re-export,
a lost test def and a changed non-source module; asserts the exact report
and the advisory/strict exit codes.
Two conflicts, both on the stderr-merge fix: managed_uv.py keeps the
fix on upstream's compact formatting; main.py taken from upstream (the
refactor moved _run_install_with_heartbeat to main_install_repair.py)
and the fix ported to the moved helper, which had reverted to a bare
subprocess.run.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
Review findings on #102117 (independent reviewer + itsflownium):
* hermes_cli.kanban_db.connect / connect_closing pointed at hermes_cli.projects_db (different DB, no
board= parameter). The compat generator ranked candidate homes by path proximity when a name is
defined in several modules. Now it requires shape compatibility with the BASE definition (same
literal for constants, superset of parameter names for defs) and prefers the facade's own
<stem>_* sibling. Same class fixed for tools.tts_tool.DEFAULT_XAI_BASE_URL (-> tts_tool_providers),
and 17 constants/defs that had been pointed at same-named strangers (Matrix MAX_MESSAGE_LENGTH ->
Signal's 8000, tts MAX_TEXT_LENGTH -> BlueBubbles', honcho/retaindb/supermemory *_SCHEMA -> another
plugin's schema, ...) are now restored from BASE verbatim instead.
* send_yuanbao_direct (restored-def): body called adapter._outbound.send_direct, which HEAD moved to
the sender; rewritten to adapter._outbound.sender.send_direct.
* COMPAT_MANIFEST.md states the scope explicitly: public top-level names only; private names and
test monkeypatch seams are not preserved.
* scripts/check_subprocess_stdin.py: _splat_carries_stdin looked 30 lines ahead in the file text
and was satisfied by an unrelated later stdin=; it now finds the splatted name's definition via AST
and requires stdin inside that expression/body.
Tests: tests/test_compat_manifest_targets.py (pointer identity vs the facade's sibling; kanban
connect(board=) opens a Kanban DB, not projects.db; both FAIL on the previous layer),
test_subprocess_stdin_guard gains the false-negative probe, and the MoA -Q quiet-output contract
tests are back (tests/agent/test_moa_quiet_reference_output.py) against build_moa_facade.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
The harness asserted the driver still throws 'not implemented yet' for
the desktop-installer@latest update route; that arm is implemented now
(Invoke-PhaseInstallGui -Mode "update"), so the check failed against
its own tree. It asserts the implemented contract instead.
The 256-job cap wording in the workflow comment and README now states
the scope GitHub applies it at: each per-OS matrix separately, not the
combined leg count. At the 10-tag bound the largest matrix is windows
at 180.
e2e-screen-record comment: hhttps -> https.
^(10|[1-9])$ replaces the two-step guard: the [0-9]+ regex accepted
leading zeros that bash arithmetic then read as octal (010 passed as 8,
08 errored). README cost figures corrected to the generator's real
expansion: 41 legs/tag, 82 at the default 2 tags, update route 8/tag,
first matrix overflow at 15 tags (270 windows entries).
tag-count now reaches the shell via the environment, validated to 1-10
(an apostrophe in the raw interpolation could terminate quoting; above
~14 tags the expansion exceeds GitHub's 256-job matrix limit).
Result-chart cell ranking matches on the leading token: rendered
success/failure cells carry artifact links, so whole-cell indexOf
ranked them -1 and any skip in the map beat a real outcome.
README documents per-run cost, route slice sizes, the tag-count bound,
and a warning against running the GUI drivers outside a disposable VM.
One conflict: upstream 6e7c7c7da9 replaced bot-mode-closed-chat-stays-closed.spec.ts with bot-mode-row-click-mirrors-registry.spec.ts while our side had rewired its mock-server import. Kept upstream's replacement and rewired the three new specs importing ./mock-server to the consolidated tests-js copy (symbols verified present).
`Start-UiServer` printed the -SelfTestUi URL (and opened the browser window)
as soon as the TcpListener was bound, but the runspace that answers /progress
starts asynchronously — BeginInvoke returns before the pipeline is open and
the script block is JIT'd, which is seconds on a loaded runner. The kernel
accepted connections into the backlog during that gap and nobody answered
them. The self-test hit it three times (#90371 and two follow-ups each
widened a timeout instead of removing the race) and it just failed an
unrelated hermes_state.py PR (run 33591547099, two 5s stale-backlog
timeouts = red).
- windows.ps1: readiness handshake after BeginInvoke — one /progress
round-trip must succeed (≤15s) before the server is returned; on failure
tear the listener down and continue without UI. The URL now means
"serving", not "bound". Also fixes the browser opening to a page that never
loads on a slow machine.
- test: 1s per-attempt probe timeout so a single dead backlog socket cannot
consume half the readiness budget.
- CI: new `desktop_updater` classifier lane. tests/test_desktop_update_windows_*.py
spawn the real PowerShell script; the Windows-only job now runs them only
when scripts/desktop-update/**, the Electron updater launcher, conftest,
pyproject, or those tests change (push/dispatch fail open). A PR that
never touched that surface cannot be failed by its process timing.
The `${{ false && (...) }}` if-expression on the reusable-workflow
e2e-desktop job made GitHub's workflow parser fail at startup (annotation:
"An unexpected error has occurred"), so every ci.yaml run on main and
every PR dispatched 0 jobs. Discriminator on this branch: pre-24f5a60
blob → 26 jobs; paren-relocated && form → 0; bare if: false → 25.
Lane stays disabled; comment documents the re-enable line.
Linux is case-sensitive; Windows and macOS are not. Two tracked paths
differing only by case (README.md vs readme.md, src/Foo.py vs SRC/foo.py)
land fine on Linux and silently break every clone on a case-insensitive
host — the filesystem holds one, so checkout fails or whichever wins
clobbers the other. Git won't stop the pair from landing; it only warns
at checkout time on a case-insensitive FS. This is the enforcement point.
Adds scripts/check-case-collisions.py (index scan keyed on casefolded
full paths) + an unconditional workflow_call job wired into ci.yaml and
the all-checks-pass gate — unconditional because a collision can ship in
any kind of PR (docs, JS, config), not just Python, so gating on a
language lane would be the same passive-rule trap the infographic check
closes. Tests in tests/scripts/test_case_collision_check.py build
collisions via git update-index --cacheinfo so they run on
case-insensitive filesystems too.
The lane was disabled Aug 2 2026 (#76627) because the mock-backend
Electron window never got a title after the Aug 1 engines/npm churn
(#76499/#76562/#76575), failing every PR identically. #99671 fixed the
root cause: per-platform/layout Electron binary resolution in the e2e
harness (apps/desktop/e2e/electron-binary.ts). The suite is green again
on Node 26 + npm 12 — delete the temporary `false &&` guard and update
the stale comment block.
Fixes#76627
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
one-line compat forwarder) and the new implementation already drains
both pipes asynchronously with bounded abandonment, which supersedes
this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
addition inside the consolidated file; rewired the five upstream specs
still importing './mock-server' to the consolidated path (export sets
verified identical) and dropped the superseded apps/desktop/e2e copy.
The slowest green leg ever recorded is 29 minutes; every cap hit in the
suite's history was a hang, never work. Caps were linux 75 / macos 120 /
windows 240, so a wedged leg burned up to 4 hours of runner time to
report what its log showed in the first minutes. 60 minutes covers the
slowest leg plus cold-cache variance, and every driver-internal bound
(dmg install 45m, AHK 50m, updater wait) still fires before the job cap
in any single-hang scenario, keeping failure diagnostics specific.
The detached-updater wait drops 90m -> 35m on the same evidence: a
working updater finishes far inside 35m; a wedged one never finishes at
any bound, and the longer wait only delayed the report by an hour.
The workflow input descriptions and the skips README still declared
open-app-update and the Setup.exe re-run as driver TODOs; both run now.
Skips have exactly two causes and the prose names them: no OS entry
point for the pair, or the starting release predates the surface. The
chart's TODO label itself stays until the n/a relabel lands with the
known-broken-OLD gate work.
The last declared TODO: a user whose install is stale re-downloads
Hermes-Setup.exe and clicks Install over the existing install, the GUI
twin of re-running the one-liner. Windows shows the full installer UI on
a re-run (the already-installed fast path is macOS-only), so the existing
AHK install drive applies unchanged; install.ps1's repository stage
fetches the existing checkout forward to what main serves, now HEAD.
Invoke-PhaseInstallGui gains an update mode instead of a parallel copy:
the phase label, proof dir, and expected-sha assertion become parameters,
and the update-is-available assert stays install-only. The bootstrap log
rotates before the re-run so the AHK's completion fallback cannot match
the install phase's old completion line.
A dmg user can also update from the terminal (hermes update) or by
re-running the install one-liner; the dmg arm only routed the two
app-button methods, leaving three declared cells as permanent TODOs.
Port the POSIX driver's method blocks into the macos driver's update
phase (hermes-update with the --yes probe, installer re-run with per-ref
flag probing, the +desktop built-app assert) and open the workflow gate.
The dmg bootstrap driver also learns to recover instead of waiting out
its bound when a stage fails: the bootstrap parks on an error screen
with a Retry button (seen live: HTTP 429 downloading install.sh under
full-matrix runner load), so the driver watches the bootstrap log for
real failure shapes, requires two consecutive error probes before
retargeting the click at Retry's measured position, unlatches when the
log goes healthy, and gives up with the true cause after three retries.
Driver rule learned three times in this suite (lsof +D, git show and
find piped to grep -q): under set -euo pipefail, never feed grep -q
from a pipe... grep exits at first match, the producer takes SIGPIPE,
and a TRUE condition reads as failure. Capture to a variable or test
paths directly.
Verified end to end: run 33506177130, macos slice 13/13 green.
Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):
- Gate every scope-path branch on a new _IS_LINUX constant instead of
'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
never touches systemd code — no probe subprocess, no scope argv,
byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
legacy argv byte-identical, no unit recorded) and probe-returns-False
off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
wired into the on-demand windows-venv-e2e lane): real spawn_local on
windows-latest asserting jobs run exactly as before — spawned, output
captured, exit code correct, systemd path never reached even under
faked gateway identity.
Refs #70716, #71378.
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
Follow-ups on top of the salvaged commits from PR #87111 (@HexLab98) and
PR #87265 (@JoaoMarcos44):
- keep main's #92991 stall watchdog (150s progress-based) as the single
steady-state liveness probe instead of adding a second overlapping one
- orphaned-client aclose() cleanup uses the wall-clock thread deadline and
is tracked in _background_tasks so a wedged close can neither hang nor
leak one task per reconnect attempt (from #87265's review findings)
- merge #87265's no-keepalive getUpdates pool (max_keepalive_connections=0)
with #87111's TCP-keepalive socket options on all transports
- add tests/gateway/test_telegram_closewait_windows_live.py: live probes
against a real half-closing HTTP server, skipif non-win32, wired into
the on-demand windows-venv-e2e lane (wine2e/**)
Real windows-latest coverage for the #88136 salvage: probes drive the
actual _git_is_trampoline/_locate_real_git/_ensure_non_trampoline_git
helpers against the runner's genuine Git-for-Windows install plus a real
fork-bomb-guard trampoline stand-in. Wired into the on-demand
windows-venv-e2e lane (wine2e/** pushes only).
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):
- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
false-positives), and an explicit start-time expectation is now honored
on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
ProcessRegistry._terminate_host_pid (previously unverified), and the
session-close path now runs the same daemon identity verification as
the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
probes (real spawned processes, real psutil ancestry) wired into the
on-demand windows-latest wine2e lane.
Fixes#98814Fixes#89614
Hermes-Setup is a Tauri app that boots to a setup-choice screen and waits
for a click on Install Hermes before any install work starts; run bare it
blocked until the 120-minute job cap (both dmg legs, every run). Launch it
in the background with the driver's env, read the window geometry via
System Events (position and size need no assistive grant), post a real
CGEvent click with cliclick at the button's measured position (65% of
window height; System Events' own click needs assistive access the runners
deny), then wait for the full install to land: checkout, venv console
script, AND the built Hermes.app, since the cleanup trap would otherwise
kill the installer before its desktop-build stage. Bounded at 45 minutes
with desktop screenshots on every phase and failure. Adds a macos-desktop
dispatch route so this arm iterates without the full matrix.
Verified end to end: run 33407401698, both dmg legs green (first ever),
macos slice 10/10.