Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The Python suite and the JS checks were split into many small jobs to make that size usable. Each split job repeated the full setup. In most of the JS jobs the repeated setup cost more than the work. The work lanes move to larger runners. Then the splits that existed only to make small runners usable go away. Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores clear the floor that the slowest single test file sets, which is about 82s. A second slice divides work that is already at that floor, and adds a second setup. Duration data from run 32522943054 gives the numbers behind this: 3178 files, 11645s in series. The worker count is explicit, because `run_tests.sh` defaults to twice the core count. A later commit sets it from a measurement on this hardware. JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to spread about 612s of work. One larger runner installs one time. The three UI shard scripts and `run-ui-shard.mjs` are therefore removed, because the unsharded `test:ui` covers the same tests. The unit of parallel work inside that job is a CHECK, and not a workspace. apps/desktop is most of the payload, and its own `check` is a serial && chain. A spread across workspaces alone therefore leaves that chain as the long pole. A package that declares `check:*` sub-scripts gives one unit for each sub-script. That is the same selection rule the matrix used. The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same sequence runs on a laptop. It runs 11 units together, buffers the output of each one, and fails at the end with the full list. Children that share one stdout interleave their lines and make a failure hard to read. `npm run --ws check` stops at the first workspace that fails. `check:test:plugins` joins the desktop `check` script. The matrix prefers `check:*` sub-scripts over the plain `check` script, so `check:test:plugins` ran only as its own leg. Without this change the merge drops that suite and the job stays green. node_modules is cached on the lockfile, and `npm ci` is skipped on an exact hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball cache, which leaves the extract and the postinstalls to pay again. The arm64 image build stays on a native arm64 runner. A build of linux/arm64 on an x64 host uses emulation. The docker test lane caps its workers at the core count. Each of those tests drives a container, so the docker daemon sets the limit and not the processor. `.github/actionlint.yaml` declares the runner labels. actionlint knows the GitHub-hosted labels only, and an undeclared label reads as an error that hides the real findings. The `detect` job checks out one file through a sparse checkout, and its timeout drops to 1 minute. It reads `scripts/ci/classify_changes.py` and nothing else. Verification: - actionlint reports 9 findings across all workflows. An unmodified HEAD with the same config reports the same 9. This change adds none. - A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and `ubuntu-latest-32-arm-cores`. - Every changed workflow parses, and `name` parses as a string. - A replay of the `save-durations` merge step against a three-artifact layout returns all 3178 entries. - An expansion of the npm script graph gives the same leaf commands for the parallel units and for a plain `npm run check`, in both directions. Against the 13-leg matrix the count is 13 to 11, and the whole difference is the three UI shards that collapse into one unsharded `check:test:ui`. - `--list` reports the 11 units, and a full local run completes and reports the time of each unit. - The runner labels cannot be verified here. The first real run is the test.
368 lines
15 KiB
YAML
368 lines
15 KiB
YAML
name: CI
|
|
|
|
# Orchestrator workflow. Runs ``detect-changes`` once, then conditionally
|
|
# calls the sub-workflows that a PR can actually affect. A final
|
|
# ``all-checks-pass`` gate job aggregates results so branch protection only
|
|
# needs to require a single check.
|
|
#
|
|
# Sub-workflows are triggered via ``workflow_call`` and keep their own job
|
|
# definitions, matrices, and concurrency settings. They no longer have
|
|
# ``push:`` / ``pull_request:`` triggers of their own — everything flows
|
|
# through this file.
|
|
#
|
|
# SECURITY: this workflow runs PR-controlled actions, workflows, and code.
|
|
# Do not add ``secrets: inherit`` or GitHub App credentials here. Trusted
|
|
# main-only automation uses protected environments in its own workflows.
|
|
|
|
on:
|
|
pull_request:
|
|
push:
|
|
branches: [main]
|
|
|
|
permissions:
|
|
contents: read
|
|
pull-requests: write # needed by lint (PR comment) + supply-chain review_status
|
|
actions: read # needed by osv-scanner (SARIF upload)
|
|
security-events: write # needed by osv-scanner (SARIF upload)
|
|
|
|
concurrency:
|
|
group: ci-${{ github.ref }}
|
|
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
|
|
|
|
jobs:
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
# detect: run the classifier once. Every downstream job reads its outputs
|
|
# to decide whether to run. On push/dispatch the classifier fails open
|
|
# (all lanes true) so post-merge validation is never weakened.
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
detect:
|
|
name: Detect affected areas
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 1
|
|
outputs:
|
|
python: ${{ steps.classify.outputs.python }}
|
|
python_prod: ${{ steps.classify.outputs.python_prod }}
|
|
frontend: ${{ steps.classify.outputs.frontend }}
|
|
site: ${{ steps.classify.outputs.site }}
|
|
scan: ${{ steps.classify.outputs.scan }}
|
|
deps: ${{ steps.classify.outputs.deps }}
|
|
uv_lock: ${{ steps.classify.outputs.uv_lock }}
|
|
npm_lock: ${{ steps.classify.outputs.npm_lock }}
|
|
installer: ${{ steps.classify.outputs.installer }}
|
|
rust: ${{ steps.classify.outputs.rust }}
|
|
docker_meta: ${{ steps.classify.outputs.docker_meta }}
|
|
mcp_catalog: ${{ steps.classify.outputs.mcp_catalog }}
|
|
ci_review: ${{ steps.classify.outputs.ci_review }}
|
|
ci_review_files: ${{ steps.classify.outputs.ci_review_files }}
|
|
event_name: ${{ github.event_name }}
|
|
steps:
|
|
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
- name: Detect affected areas
|
|
id: classify
|
|
uses: ./.github/actions/detect-changes
|
|
with:
|
|
sparse-checkout: scripts/ci/classify_changes.py
|
|
sparse-checkout-cone-mode: false
|
|
github-token: ${{ github.token }}
|
|
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
# Lane-gated sub-workflows. Each runs in parallel after detect finishes.
|
|
# Skipped workflows (if condition is false) don't spin up runners.
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
tests:
|
|
name: Python tests
|
|
needs: detect
|
|
if: needs.detect.outputs.python == 'true'
|
|
uses: ./.github/workflows/tests.yml
|
|
|
|
# macOS + Windows lanes. The main `tests` lane above is Linux-only, and
|
|
# the OS-marked tests it collects are skipped there by design (see the
|
|
# `_OS_MARKS` comment in tests/conftest.py) — this is where they run.
|
|
# Same `python` lane gate: if no Python changed, neither runs.
|
|
tests-os:
|
|
name: OS-specific tests
|
|
needs: detect
|
|
if: needs.detect.outputs.python == 'true'
|
|
uses: ./.github/workflows/tests-os.yml
|
|
|
|
lint:
|
|
name: Python lints
|
|
needs: detect
|
|
if: needs.detect.outputs.python == 'true'
|
|
uses: ./.github/workflows/lint.yml
|
|
with:
|
|
event_name: ${{ needs.detect.outputs.event_name }}
|
|
|
|
js-tests:
|
|
name: JS & TS checks
|
|
needs: detect
|
|
if: needs.detect.outputs.frontend == 'true'
|
|
uses: ./.github/workflows/js-tests.yml
|
|
|
|
installer-tests:
|
|
name: Installer tests
|
|
needs: detect
|
|
# Windows-only, and only for PRs that touch install.ps1 or its tests.
|
|
if: needs.detect.outputs.installer == 'true'
|
|
uses: ./.github/workflows/installer-tests.yml
|
|
|
|
rust-tests:
|
|
name: Rust tests
|
|
needs: detect
|
|
# Only for PRs that touch a Rust crate. `.rs` is under apps/, so these
|
|
# changes used to run the TypeScript matrix and nothing that compiles them.
|
|
if: needs.detect.outputs.rust == 'true'
|
|
uses: ./.github/workflows/rust-tests.yml
|
|
|
|
e2e-desktop:
|
|
name: Desktop E2E
|
|
needs: detect
|
|
# python_prod (not python): the Playwright suite exercises the built app
|
|
# + `hermes serve` backend, which never import anything under tests/.
|
|
# Tests-only PRs (~17% of commits) skip this 5-minute job — the longest
|
|
# single job in the workflow — while still running the full pytest lanes.
|
|
#
|
|
# ⛔ TEMPORARILY DISABLED (Aug 2, 2026, Teknium) — the suite is red on
|
|
# every PR and on main itself since the Aug 1 night engines/npm churn
|
|
# (#76499 → #76562 → #76575): the mock-backend Electron window never
|
|
# gets a title, so boot/chat/setup/interim specs all fail identically
|
|
# regardless of the PR's diff (verified on #76573 and the docs-only
|
|
# #76582). Tracking issue: #76627 (assigned: Ari). To re-enable,
|
|
# delete the `false &&` below — nothing else changed.
|
|
if: ${{ false && (needs.detect.outputs.python_prod == 'true' || needs.detect.outputs.frontend == 'true') }}
|
|
uses: ./.github/workflows/e2e-desktop.yml
|
|
|
|
docs-site:
|
|
name: Docs Site
|
|
needs: detect
|
|
if: needs.detect.outputs.site == 'true'
|
|
uses: ./.github/workflows/docs-site-checks.yml
|
|
|
|
history-check:
|
|
name: Deny unrelated histories
|
|
needs: detect
|
|
if: needs.detect.outputs.event_name == 'pull_request'
|
|
uses: ./.github/workflows/history-check.yml
|
|
|
|
contributor-check:
|
|
name: Check contributors
|
|
needs: detect
|
|
if: needs.detect.outputs.python == 'true'
|
|
uses: ./.github/workflows/contributor-check.yml
|
|
|
|
uv-lockfile:
|
|
name: Check uv.lock
|
|
needs: detect
|
|
# Gated: `uv lock --check` re-resolves the whole dependency graph against
|
|
# PyPI, so on every PR it spent a network round-trip — and, on a registry
|
|
# blip, a blocking red X — for diffs that cannot desync the lockfile
|
|
# (docs, frontend, prose). Only pyproject.toml / uv.lock can. A
|
|
# `.github/` change still forces it on via the classifier's fail-open.
|
|
if: needs.detect.outputs.uv_lock == 'true'
|
|
uses: ./.github/workflows/uv-lockfile-check.yml
|
|
|
|
infographic-check:
|
|
name: Check no committed infographics
|
|
needs: detect
|
|
uses: ./.github/workflows/infographic-check.yml
|
|
|
|
lockfile-diff:
|
|
name: package-lock.json diff
|
|
needs: detect
|
|
if: needs.detect.outputs.event_name == 'pull_request' && needs.detect.outputs.npm_lock == 'true'
|
|
uses: ./.github/workflows/lockfile-diff.yml
|
|
|
|
docker-lint:
|
|
name: Lint Docker scripts
|
|
needs: detect
|
|
if: needs.detect.outputs.docker_meta == 'true'
|
|
uses: ./.github/workflows/docker-lint.yml
|
|
|
|
supply-chain:
|
|
name: Supply-chain scan
|
|
needs: detect
|
|
if: needs.detect.outputs.event_name == 'pull_request' && (needs.detect.outputs.scan == 'true' || needs.detect.outputs.deps == 'true')
|
|
uses: ./.github/workflows/supply-chain-audit.yml
|
|
with:
|
|
event_name: ${{ needs.detect.outputs.event_name }}
|
|
scan: ${{ needs.detect.outputs.scan == 'true' }}
|
|
deps: ${{ needs.detect.outputs.deps == 'true' }}
|
|
|
|
review-labels:
|
|
name: Review label gate
|
|
needs: [detect, supply-chain]
|
|
if: always() && needs.detect.outputs.event_name == 'pull_request' && (needs.detect.outputs.ci_review == 'true' || needs.detect.outputs.mcp_catalog == 'true' || needs.supply-chain.outputs.critical_findings == 'true')
|
|
uses: ./.github/workflows/review-labels.yml
|
|
with:
|
|
ci_review: ${{ needs.detect.outputs.ci_review == 'true' }}
|
|
ci_review_files: ${{ needs.detect.outputs.ci_review_files }}
|
|
mcp_catalog: ${{ needs.detect.outputs.mcp_catalog == 'true' }}
|
|
supply_chain: ${{ needs.supply-chain.outputs.critical_findings == 'true' }}
|
|
|
|
osv-scanner:
|
|
name: OSV scan
|
|
uses: ./.github/workflows/osv-scanner.yml
|
|
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
# Gate: runs after everything. ``if: always()`` ensures it reports a
|
|
# status even when some deps were skipped. Only actual ``failure``
|
|
# results cause it to fail; ``skipped`` is treated as success.
|
|
#
|
|
# Branch protection should require ONLY this check.
|
|
#
|
|
# Outputs ``needs-json`` — a compact ``{job_name: result}`` dict — so
|
|
# the live comment poller can list failed jobs in the PR comment.
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
all-checks-pass:
|
|
name: All required checks pass
|
|
needs:
|
|
- detect
|
|
- tests
|
|
- tests-os
|
|
- lint
|
|
- js-tests
|
|
- installer-tests
|
|
- rust-tests
|
|
- e2e-desktop
|
|
- docs-site
|
|
- history-check
|
|
- contributor-check
|
|
- uv-lockfile
|
|
- lockfile-diff
|
|
- docker-lint
|
|
- supply-chain
|
|
- review-labels
|
|
- osv-scanner
|
|
# The image build runs in its own workflow (docker.yml) and reports
|
|
# its own check. It was never required here, because it is too slow
|
|
# to block a merge. A separate run also stops it from holding this
|
|
# run open. That is what blocked ``gh run rerun``.
|
|
if: always()
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
outputs:
|
|
needs-json: ${{ steps.evaluate.outputs.needs-json }}
|
|
steps:
|
|
- name: Evaluate job results
|
|
id: evaluate
|
|
env:
|
|
NEEDS: ${{ toJSON(needs) }}
|
|
run: |
|
|
echo "$NEEDS" | python3 -c "
|
|
import json, sys
|
|
needs = json.load(sys.stdin)
|
|
# Emit compact {job_name: result} for the comment assembler.
|
|
compact = {name: info['result'] for name, info in needs.items()}
|
|
print(f'needs-json={json.dumps(compact)}')
|
|
with open('$GITHUB_OUTPUT', 'a') as f:
|
|
f.write(f'needs-json={json.dumps(compact)}\n')
|
|
failed = [name for name, info in needs.items() if info['result'] == 'failure']
|
|
for name, info in sorted(needs.items()):
|
|
result = info['result']
|
|
icon = '✅' if result in ('success', 'skipped') else '❌'
|
|
print(f'{icon} {name}: {result}')
|
|
if failed:
|
|
print(f'::error::{len(failed)} job(s) failed: {\", \".join(failed)}')
|
|
sys.exit(1)
|
|
print('All checks passed (or were skipped)')
|
|
"
|
|
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
# CI timing report: collect per-job/step durations from the GitHub API,
|
|
# cache them on main (as a baseline), and on PRs generate an HTML diff
|
|
# report with a gantt chart + per-step breakdown. The report is uploaded
|
|
# as an artifact and a markdown summary is written to $GITHUB_STEP_SUMMARY.
|
|
#
|
|
# The live comment poller dynamically fetches all review-status-* artifacts
|
|
# across the orchestrator and sub-workflow runs every cycle, so its link
|
|
# points straight at that report.
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
ci-timings:
|
|
name: CI timing report
|
|
needs: [all-checks-pass]
|
|
if: always()
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
steps:
|
|
- name: Checkout code
|
|
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
|
|
- name: Restore baseline cache (PR only)
|
|
if: github.event_name == 'pull_request'
|
|
uses: actions/cache/restore@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
|
|
with:
|
|
path: ci-timings-baseline.json
|
|
# Prefix-match: exact key will never hit (run_id differs), so
|
|
# restore-keys finds the most recent baseline from main.
|
|
key: ci-timings-baseline-never-exact
|
|
restore-keys: |
|
|
ci-timings-baseline-
|
|
|
|
- name: Collect timings and generate report
|
|
env:
|
|
GITHUB_TOKEN: ${{ github.token }}
|
|
run: |
|
|
python3 scripts/ci/timings_report.py \
|
|
--baseline ci-timings-baseline.json \
|
|
--output ci-timings-report.html \
|
|
--json-out ci-timings.json \
|
|
--summary-out ci-timings-summary.md
|
|
|
|
- name: Upload HTML report
|
|
# Advisory report — artifact-service blips must not fail the job.
|
|
continue-on-error: true
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
id: ci-timings-html
|
|
with:
|
|
name: ci-timings-report
|
|
path: ci-timings-report.html
|
|
retention-days: 14
|
|
|
|
- name: Build linked review status
|
|
if: hashFiles('ci-timings.json') != ''
|
|
env:
|
|
CI_TIMINGS_REPORT_URL: ${{ steps.ci-timings-html.outputs.artifact-url }}
|
|
run: |
|
|
python3 scripts/ci/timings_report.py \
|
|
--from-json ci-timings.json \
|
|
--baseline ci-timings-baseline.json \
|
|
--review-status-out review-status.json \
|
|
--review-status-only
|
|
|
|
- name: Upload review status
|
|
if: hashFiles('review-status.json') != ''
|
|
continue-on-error: true
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: review-status-ci-timings
|
|
path: review-status.json
|
|
retention-days: 14
|
|
|
|
- name: Output summary
|
|
env:
|
|
REPORT_URL: ${{ steps.ci-timings-html.outputs.artifact-url}}
|
|
run: |
|
|
{
|
|
echo "# CI Timing report"
|
|
echo "[View the full interactive report]($REPORT_URL)"
|
|
} >> "$GITHUB_STEP_SUMMARY"
|
|
cat ci-timings-summary.md >> "$GITHUB_STEP_SUMMARY"
|
|
|
|
- name: Save baseline cache (main only)
|
|
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
|
|
run: |
|
|
# Degraded runs (API rate-limited) produce no ci-timings.json —
|
|
# skip rather than fail, and never cache an empty baseline.
|
|
if [ -f ci-timings.json ]; then
|
|
cp ci-timings.json ci-timings-baseline.json
|
|
else
|
|
echo "No timings JSON this run — skipping baseline update"
|
|
fi
|
|
|
|
- name: Upload baseline to cache (main only)
|
|
if: github.event_name == 'push' && github.ref == 'refs/heads/main' && hashFiles('ci-timings-baseline.json') != ''
|
|
uses: actions/cache/save@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
|
|
with:
|
|
path: ci-timings-baseline.json
|
|
key: ci-timings-baseline-${{ github.run_id }}
|