Keep the minimal sandbox command's existing app/data identity while sharing
its parser, seeding, execution and cleanup with the standard sandbox.
Remove the unused SDK adapter copy and protocol shadow from the environment
facade, plus unconsumed checkout-updater dependencies and their dead helper.
Validation: 72 focused Python tests across sandbox and adapter/backend
paths, 5 checkout-updater tests, Electron typecheck and ESLint passed.
Ruff, touched-file Windows checks and the TS ratchet pass. Sandbox tests
also passed against the original scripts before consolidation.
Keep desktop rollback in staged publication instead of restoring backups
inside a candidate directory that will be discarded.
Use the npm tag endpoint, let the downloader own archive verification,
and ensure CI toolchain roots without repeating dependency verification.
Remove an unused resolver input and replace overlapping tests with
fault-proven lifecycle coverage.
Verified with the scoped Python gate and before-pack tests. Native
Windows/macOS update execution and the full suite were not run.
Preserve upstream fixes without restoring retired dependency installers.
Run configured-feature checks in the selected build interpreter. Reuse a
supported base Python during bootstrap, and preserve durable backup media.
Refresh the dependency lock through PM. Keep the frozen historical import
surface unchanged. Adapt incoming native tests to the platform markers.
Verification: the incoming 86-file pass found two fixture mismatches;
both passed after correction. Targeted PM/update/compatibility checks,
Electron and renderer typechecks, and desktop tests passed.
Native Windows/macOS update journeys and the full suite remain unrun.
Run historical updater completion in a fresh interpreter so cached imports
cannot revive retired dependency installers. Share Git and ZIP completion,
carry receipt and recovery state, and preserve child exit status.
Route plugin admission, binary acquisition, desktop launch and build paths
through PM. Replace redundant helpers and tests with real worker, package,
publication and launch checks. Keep the shipped compatibility surface fixed.
Targeted Python and desktop checks pass. Native update journeys and fresh
production image qualification remain pending. This is a checkpoint before
those acceptance runs.
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.
Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.
Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.
Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.
Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.
Give native payload dependency builds two hours without changing the normal install timeout. Allow three hours for the standalone PM bundle job so setup, cache saves, and smoke tests fit around the build.
Verified seven focused tests, Ruff, and actionlint. Full CI builds were not rerun.
Resolve the default cache before isolating HOME so payload builds use the directory that CI restores and saves. Remove the standalone bundle workflow dependency on an unset cache variable.
Verified offline wheel reuse in isolated children, 20 focused tests, Ruff, and actionlint. Full native release builds and the separate source-build timeout remain unverified.
Commit builds have no update channel, even when installed through dpkg. Read the installed stamp before checking the refusal message and require exit code 2 for both artifact kinds.
Verified real CLI refusal paths, negative controls, and related tests: 15 passed. Ruff and type checks passed. Full deb validation was not rerun.
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute
job timeout, so the live skills-index.json has been frozen at that date and every
`hermes skills search` fell through to live GitHub API calls (~500 inspect requests per
cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been
appending to #66616 four times a day since.
Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to
fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k
skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it.
- enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder
without an owner (the link is a nicety; the index is not).
- build_skills_index.py passes an 8-minute budget.
- skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path
(clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min).
Reuse a matching ARM64 Visual Studio environment instead of growing its paths on each initialization. Preserve Rust and Cargo homes before isolating bundle state. Hand the private Termux assembly directory to its container user and restore host ownership afterward.
Verified targeted tests, real ARM64 SDK/OpenSSL execution, Rust compilation, and container facts publication. Full release packages were not rebuilt. The Windows x64 source-build timeout and the unchanged Termux linkage fixture failure remain unresolved.
Canary and commit builds must not replace stable or share its desktop
state. Package names alone are insufficient because Electron reads the
product name before main initializes its paths. Pin nonstable userData
before the first lookup, and keep the packaged identity independent of
runtime build variables.
Keep release artifact filenames unchanged. Qualify payload CLI names,
route each nonstable MSIX alias to its own entrypoint, and copy the
immutable desktop provenance into the embedded Python checkout. Only
stable releases can use the official Store identity.
Targeted validation: 75 JavaScript tests passed, 2 platform skips;
15 Python tests passed with file retries disabled. Native Windows SDK
manifest proof is tracked separately. Full app install, signing and macOS
launch validation are not claimed.
Canary uses yellow tiles. Commit builds use red tiles and the first seven
characters of HERMES_BUILD_COMMIT. Vector glyphs need no host fonts.
Only desktop targets use these colors. Shared branding remains unchanged.
PM-generated pixels preserve the artwork and native alpha masks. The macOS
content bounds remain (100, 100, 924, 924) on the 1024 canvas. Stable output
matches the pre-change baseline byte for byte.
Validation: 13 focused Python tests and 7 Node tests pass. Stable, canary,
and commit generation each wrote 35 targets and passed structural checks.
The full suite and native installation acceptance were not run.
The minted launchers wired venv site-packages onto sys.path with a raw
insert (win32 wrapper) / PYTHONPATH (posix), neither of which runs .pth
files. pywin32.pth is load-bearing on Windows: it puts win32\lib on
sys.path, which is what makes 'import pywintypes' resolve — without it
portalocker's Win32Locker dies and concurrent-log-handler silently drops
every file-log record on Windows bundles.
* launcher_wrapper.py: site.addsitedir() for the site entry (repo first,
site directly after, .pth dirs last)
* launchers.py posix: same via HERMES_SITE env in the -c bootstrap
* pm/environment.py: prune_site_pth() drops _virtualenv.pth and the
__editable__ pointer (build-machine path) that must never run in a
sealed payload
* python_env.py: run the prune after every environment build
An inherited HERMES_HOME can defeat a test bundle's data-directory suffix.
Older Windows installers also persisted that variable in the user registry.
This gives the app fresh UI state while its backend reads existing sessions.
Add --bundle-unset NAME, encoded as null in the existing bundle environment
object. Apply each clear as an explicit empty value before module startup.
Do not restore an explicitly empty HERMES_HOME from the Windows registry.
Ordinary defaults still preserve runtime overrides.
Verified the release parser, builder handoff, compiled startup ordering,
registry opt-out, and child environment with focused regression tests.
A native Windows probe passed with an inherited home. No MSIX was rebuilt.
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three
model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces.
Later non-Wisdom work on shared files (guest onboarding i18n, dashboard
startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call
sites and config were stripped from those files.
ntpath.isreserved was added in Python 3.13, but release workflow legs
(commit-builds-summary, builds-table, builds-pending) run bare python3 on
ubuntu-24.04, whose system Python is 3.12. render-builds-table.py crashed
with AttributeError before rendering the expected-binary matrix.
Port the CPython ntpath reserved-name semantics (device stems incl.
superscript COM/LPT forms, trailing dot/space per component) into
_is_windows_reserved() in scripts/releases/r2.py so the release transport
stays self-contained on whatever python3 the runner provides. Verified
byte-parity against real ntpath.isreserved on a 239-case corpus.
Merge ethie/shared-product-builders with the CI dependency cache and native Windows setup work. Preserve UTF-8 diagnostics in the shared Python environment runner. Pass a persistent cache through isolated native staging and PM-runtime construction. Reuse one Windows prerequisite installer from source setup, native adapters, and CI, preserving Rust homes across HOME isolation.
Verified 85 targeted Python tests (5 host skips), 18 JavaScript tests, workflow validation, and scoped lint/typecheck. On native Windows ARM64, five prerequisite contracts passed and the actual shared provider reused OpenSSL, compiled its header with MSVC, and retained Rust under isolated HOME. Full signed distribution builds and live Actions cache transfer remain CI verification.
Build TUI, web, desktop UI and runnable agent products from explicit
prepared inputs. Keep dependency preparation separate from distribution
packaging, with PM and native builds sharing uv environment construction.
Docker copies compiled frontend products instead of build dependencies.
Nix retains uv2nix environments and consumes shared assembly through store
references. Native desktop and Termux use the same launcher and frontend
contracts. Preserve the independent PM runtime and source imports from
arbitrary working directories.
Keep failed frontend builds from replacing the previous product, reject
source/output overlap, and bound dependency-process output draining.
Include hermes_wisdom in the Nix wheel: real CLI smoke tests exposed its
missing package declaration on the base revision too.
Verified focused Python and JavaScript suites, Docker build/runtime checks,
Nix desktop and CLI/ACP checks, standalone TUI and packaged Electron PTY,
and real full-Chromium interaction. Native signed installers, Android device
installation and the full repository suite remain CI verification.
Activation reaches plugin discovery before the application dependencies
exist. Give PM its own locked Python project and runtime so it can install
or repair the application without importing that dependency tree.
Keep PM outside the application workspace. A shared uv workspace resolves
the application graph and cannot provide this isolation. Route mutations
through an isolated worker and preserve transaction callbacks, cancellation,
custom package registrations, and correlated receipts.
Use the same runtime builder for source installs and packaged payloads.
Keep offline wheelhouse support in that builder. Nix builds the independent
PM lock as a separate derivation. Refuse lazy-disabled bootstrap before
installing tools or dependencies.
Move first-party YAML readers and writers to ruamel. Keep the application
lock's transitive PyYAML requirements for third-party packages.
Verification:
- Focused canonical Python suite: 177 passed, 1 host-gated skip.
- Electron backend probes: 12 passed. Electron typecheck passed.
- Both uv locks, scoped lint, Bash syntax, and whitespace checks passed.
- Cold activation, corrupt-app repair, offline staging, and relocation ran.
- Built and exercised the Nix PM runtime and standalone YAML merge script.
Six broader caller test files retain the same 24 failing test IDs as an
archive of HEAD. The existing real-home guard blocks those tests before
they can exercise the affected paths. No full-suite pass is claimed.
Native Windows signing and full Bionic package execution remain unverified.
* feat(wisdom): add trusted publish and install foundation
* feat(wisdom): add private contribution loop
* feat(wisdom): add managed consumption workflows
* fix(wisdom): close cross-repository safety gaps
* fix(wisdom): align local package and lifecycle policy
* fix(wisdom): require explicit profile setup
* docs(wisdom): repin reconciled gateway head
* fix(wisdom): fence content downloads and approval receipts
* docs(wisdom): record generation-fenced downloads
* docs(wisdom): record unified delivery PR
* fix(ci): stop passing invalid classifier inputs
* docs(wisdom): remove internal requirements ledger
* feat(wisdom): localize dashboard and desktop copy
* feat(wisdom): complete local contribution and consumption UX
* style(wisdom): satisfy desktop lint
* chore(wisdom): refresh requirements pin
* test(dashboard): allow formatted profile copy
* test(wisdom): stabilize desktop interaction coverage
* fix(wisdom): surface dashboard action failures
* fix(wisdom): add repeatable Portal demo login
* feat(wisdom): add actionable skill notifications
* feat(wisdom): add notification install and update actions
* fix(wisdom): make Telegram skill alerts actionable
* fix(wisdom): always refresh demo Agent login
* feat(wisdom): embed Telegram notification actions
* fix(wisdom): preserve Telegram notifications after actions
* fix(wisdom): keep Telegram notification cards readable
* feat(wisdom): add Telegram candidate approval flow
* feat(wisdom): explain Telegram qualification reasons
* fix(wisdom): reconcile cross-surface candidate actions
* feat(telegram): add Collective Wisdom management command
* chore(wisdom): refresh Gateway contract pin
* chore(wisdom): advance Gateway contract pin
* feat(wisdom): align command UX across clients
* feat(slack): add Collective Wisdom management parity
* feat(wisdom): add security and professionalism reviews
* feat(wisdom): add first-time qualification guidance
* feat(wisdom): simplify qualification sharing choices
* feat(skills): add optional editorial metadata
* feat(wisdom): enrich legacy skill presentation
* fix(wisdom): harden review and update boundaries
* fix(wisdom): emit canonical review timestamps
* fix(wisdom): align with merged gateway and main
* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)
- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
7-day evidence builder that excludes bundled/hub/managed skills and
dismissed/handled/recently-suggested content hashes, strict pydantic
schemas for agent output with repair-or-reject, fixed copy templates
(Share / Teammate / Published / Update / Mute), idempotent retried
delivery ledger with stale-action resolution, weekly review job,
resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.
* wisdom: agent-led renderers and button action dispatcher
- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
packaging flow, Install/Update -> plan command. Never publishes/installs.
* wisdom: CLI verbs, agent_led config default, conversational catalog skill
- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
verbs, share/install flows and fixed notification templates.
* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons
- gateway housekeeping tick calls maybe_run_weekly_review with a home
channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
duration keyboard, send_wisdom_agent_recommendation rich card + fallback.
* fix(wisdom): integrate local mediation and harden model and setup boundaries
* fix(wisdom): honor authoritative recommendation policy and defer on failure
* fix(wisdom): synchronize opaque suppression and recheck delivery preferences
* feat(wisdom): route weekly selection through the session-owned assessment queue
* fix(wisdom): prepare and submit the reviewed generated share package
* feat(wisdom): separate native Share preparation from publication consent
* feat(wisdom): sync native mute choices through a leased preference outbox
* feat(wisdom): bind native mute controls to durable preference choices
* feat(wisdom): add scoped desktop and dashboard notification settings
* fix(wisdom): revalidate feed recommendations before assessment and delivery
* fix(wisdom): persist validated delivery receipts before completing notices
* feat(wisdom): add private notification claim and receipt client
* Persist Wisdom send reservations and recover delivery acknowledgements
* Route legacy Wisdom controls through current native review
* Add typed private Wisdom operation outcome client
* fix(wisdom): make agent-led advice usable in the local demo
* fix(wisdom): keep requested consent outside proactive limits
* fix(wisdom): distinguish unavailable assessments and preserve digest text
* fix(wisdom): assess ongoing usefulness beyond the current task
* fix(wisdom): restore immediate qualification sharing controls
* fix(wisdom): separate qualification review from installation advice
* fix(wisdom): collapse review checklists and simplify sharing copy
* fix(wisdom): show compact sharing progress and publication receipts
* fix(wisdom): require credential prefixes rather than matching skill names
* fix(wisdom): finish package checks before presenting sharing consent
* fix(wisdom): scan local skills before qualification cards
* fix(wisdom): update moderation results on existing sharing cards
* fix(wisdom): keep sharing review accessible from receipt cards
* fix(wisdom): align mediated review cards and collapsible checks
* fix(wisdom): clarify clean security summary wording
* fix(wisdom): normalize consent plans and add explicit recheck
* fix(wisdom): keep install and update receipts concise
* fix(wisdom): collapse assessments and deduplicate operation cards
* fix(wisdom): restore private Portal review from native cards
* fix(wisdom): sync Portal publication to original consent card
* fix(wisdom): show local skill version on sharing cards
* fix(wisdom): skip agent recommendations for self-published versions
* fix(wisdom): simplify candidate notices and local-edit recovery copy
* feat(wisdom): submit locally reviewed packages with one confirmation
* feat(wisdom): expose safe receipt and outcome sync recovery
* wisdom: onboarding notice says detect and share, names the user's own skill
Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark
Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.
* wisdom: one opener, no approval line, ask to share after the skill is shown
Product owner review of the candidate card.
- The Hermes written card now opens with the same sentence as the fixed card
("Your organisation has enabled Collective Wisdom, a feature designed to
automatically detect and share useful skills across all team members.")
instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
It is now the last line, after the skill name, description, why suggested
and the checks, and reads "Would you like to share it?" (matching the
agent led template wording).
Tests updated for the new order; proposalNotice removed from all desktop locales.
* wisdom: American spelling, organization
Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.
* wisdom: candidate card copy round 4 (owner review)
Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:
1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
card (Telegram rich card and plain fallback, legacy agent-led share
template).
3. The skill name and description are labelled: "Skill name: <name>" and
"What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
inappropriate content found)" with no per-check bullets and no "Pass";
a failed review reads "Needs a look before sharing at work (possible
inappropriate content)" and lists only the checks that flagged
something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
editorial_name, a simple one_line_description and a compelling
why_coworkers_benefit under 300 characters; "Be concise and
convincing." becomes "Be concise and compelling: the goal is that the
user wants to share it."
Tests updated for the new strings; review_text() gains direct coverage.
* wisdom: re-apply owner copy after rebase
- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice
* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors
Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.
* fix(wisdom): reconcile optional SDK tests and frontend lint
* fix(wisdom): default to agent-written notification summaries
* fix(wisdom): restore deferred install review and browse controls
* feat(wisdom): inspect installed setup with exact package provenance
* feat(wisdom): run native-approved installed setup steps with durable evidence
* fix(wisdom): recover interrupted setup with explicit native consent
* feat(wisdom): hand native installs into guided setup review
* fix(wisdom): continue requested setup with fixed notification copy
* fix(wisdom): preserve setup while waiting for a session model
* fix(wisdom): expose canonical setup review controls on desktop
* fix(wisdom): resume setup after recorded automatic updates
* fix(wisdom): make missing setup prerequisites recheckable
* chore(wisdom): align Agent with verified Gateway contract
* fix(wisdom): stop guessing team slugs in portal links
* fix(wisdom): retire pending advice on account sign-out
* fix(wisdom): cancel advice after terminal account revocation
* fix(wisdom): fence feed responses across account sign-out
* fix(wisdom): checkpoint signed-out feed before reactivation
* fix(wisdom): link proactive advice to scoped notification settings
* fix(wisdom): coalesce queued publication recommendations by version
* fix(wisdom): keep package review navigation local and deferable
* fix(wisdom): reflect installed state in discovery controls
* fix(wisdom): show exact checks before command confirmation
* chore(wisdom): pin bounded analytics privacy contract
* chore(wisdom): pin retired legacy notification contract
* feat(wisdom): review publisher usage with exact sharing copy
* fix(wisdom): align discovery and review check summaries
* fix(wisdom): show expired consent before confirmation
* fix(wisdom): require fresh review for legacy install controls
* fix(wisdom): preserve review expiry across check toggles
* fix(wisdom): retain update policy in native install reviews
* fix(wisdom): surface failed native card edits
* fix(wisdom): persist local command approval reviews
* fix(wisdom): use saved approvals for messaging commands
* test(wisdom): provide scan result in setup handoff fixture
* test(wisdom): exercise Telegram approvals with saved review state
* fix(wisdom): retain suppression policy for offline deferral
* fix(wisdom): reconsider candidates after deferred suppression expires
* fix(wisdom): bind review checks and report verified readiness separately
* fix(wisdom): persist accepted publication intent and recover exact outcomes
* fix(sync): pin UTF-8 tree ordering across writers
* chore(wisdom): pin organisation-scoped Gateway authorization
* fix(wisdom): restrict consent delivery to user-facing sessions
* chore(wisdom): refresh reviewed Gateway contract pin
* fix(wisdom): preserve kept tools in Blank Slate exclusions
* test(auth): reset anonymous fixture with a profile-scoped cache
* fix(wisdom): gate local surfaces and work on current profile entitlement
* fix(wisdom): invalidate quiet tool cache on entitlement changes
* test(wisdom): authorize local consent gateway fixtures
* fix(wisdom): keep entitlement decoding free of native crypto imports
* test(wisdom): provide local entitlement to demo CLI subprocess
* ci: leave upstream workflow unchanged in Wisdom PR
* fix(wisdom): ship package and contracts in Nix wheels
---------
Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
Termux removes old package files, so a pinned URL and hash do not keep
build inputs available. Preserve the exact bytes without changing pins.
Archive every PM HTTP artifact and the Termux runtime inputs by SHA256.
CI reads R2 first. Only a missing object permits an upstream download,
hash verification, immutable upload, and verified readback. Seed the
actual toolchain and payload stores before their consumers run.
Use the public archive as a pinned fallback in PM, bootstrap installers,
and Nix fetchers. Keep network retries bounded and report attempted URLs.
Keep publication credentials in protected CI jobs, not installed clients.
Verification:
- 283 targeted tests passed; five POSIX tests skipped on Windows.
- All 87 preserved Termux packages passed local archive miss/hit checks.
- Native ARM64 ripgrep installed through the mirror and ran successfully.
- Wheel import, workflow lint, Python lint, shell syntax, and pins passed.
Live R2 publication, POSIX tests, and Nix builds remain for native CI.
The real-byte archive checks used loopback HTTP, not the live bucket.
Full Chromium serves both headed and headless sessions. The separate
shell duplicates the browser payload and is not needed for either mode.
Remove the shell from PM and Docker. Select the managed Chromium
executable for agent-browser and the full Chromium channel for direct
Playwright callers. Route setup through PM and remove retired packages
from cached bundle stores without changing the user's tool store.
Update signing, architecture checks, launch probes and install guidance.
Leave llama packages and Docker archive cleanup unchanged.
Verification:
- Real agent-browser navigation, clicks, DOM reads and screenshots pass
in headed and headless modes with the same Chromium executable.
- The direct Playwright doctor probe passes.
- Focused Python and desktop packaging tests pass, as do six Docker
checks and both real-browser task-scroll tests.
- The built linux/amd64 image is 1.393 GB compressed, 223.6 MB smaller.
- The broader PM suite and two unrelated setup tests still fail.
Those failures reproduce on unchanged HEAD.
- Five updated eval scripts parse; their full scenarios were not run.
The builds table only ever existed inside a GitHub release body. Emit the
same rows as a tiny standalone page in R2, so a build is readable straight
from the download origin:
releases/<channel>/index.html latest stable / canary builds, every
variant, replaced by each tag run
releases/commit/<sha>/index.html every expected binary of one commit
build, built or not
scripts/render-builds-table.py keeps ONE row set per mode and renders it
into two sinks (release body markdown, page HTML), so the page can never
list different artifacts than the release. A channel page is a mutable
pointer written from a per-tag job, so it records its release tag and the
writer compares that against scripts/releases/semver.py before replacing:
re-running an older tag cannot regress a newer channel page.
Pages need two registrations to be usable: `.html` maps to
text/html; charset=utf-8 in release-content-types.json (unregistered, R2
serves the object as an octet-stream download) and to no-store in
r2.cache_control_for (the page is a pointer, not an artifact). Page keys
and public URLs come from new r2 layout helpers, shared with
`release.py --build-commit`, which now prints the commit page URL before
dispatching. No workflow change: the existing renderer jobs already carry
the R2 credentials.
Verified: 76 tests over the renderer/release/transport files, including a
loopback R2 PUT proving the page object lands as text/html with no-store.
Keep commit admission on the trusted workflow checkout and reject mixed
release inputs before loading repository code. Stage every built product
under its commit with receipt-bound summary links, never channel writes.
Build both Windows universal bundles through the existing SDK scripts.
Keep Store calendar versions separate from sideload app versions so zero-
major app versions remain packageable. Reject invalid arguments before
modifying bundles. Bind desktop and Termux versions to the source commit,
and record Termux cache provenance without labeling commits as tags.
Verification: 77 Python tests and 36 JS tests passed. Real makeappx packed
and unpacked disposable per-arch and universal packages. Seven official
workflow-expression checks, actionlint, syntax, lint and prose passed.
No signing, installed-app update, Android build, or remote dispatch ran.
Resolve pushed revisions before dispatching the default-branch workflow.
Reject release-mode flags and untrusted admission contexts. A dry run
never dispatches or creates a tag. Preserve Git's effective push URL
when choosing the GitHub repository.
Commit-build stamps check the actual checkout, including an explicit
Python --commit argument. The workflow SHA cannot replace build identity.
Direct Git argv also avoids the Windows command-shell PATH limit.
Real temporary Git CLI and stamp tests pass: 60 Python tests and 22 JS
tests, with no failures. GitHub authorization and dispatch are intercepted
at their process boundary. Workflow guards and native assembly remain
separate work. No live dispatch, signature, or package acceptance claimed.
Staging digests became stale after PE repair or payload signing.
Refresh all tool facts atomically while preserving their identity.
Windows refreshes after the batch signer. The macOS wrapper delegates
to the installed signer and refreshes after children, before the outer
app signature. Entitlements and file selection remain intact.
The real builder dispatcher caught the rejected factory shape. Tests
execute the installed signing walk with only codesign intercepted,
then compare the recorded bytes at the outer signing boundary.
Corrupt facts abort that boundary. Enabling batching fails the tests.
Verification: 80 Python tests passed with 3 host skips. 36 JavaScript
tests passed with 1 host skip. Lint, config typecheck and schema pass.
No real Apple/Azure signature, native package install or release run.
The builder's imported modules could mark an empty dependency tree as
installed. Probe every anchor in one isolated target process. Require
all anchors for a multi-module extra, and process only the selected
tree's .pth files so editable packages retain their launch behavior.
The native builder passes its staged Python explicitly and stops before
manifest publication when the inventory fails. Correct the Hindsight
and Teams anchors using the namespaces in their locked wheels.
Verified: 31 tests passed, 3 host skips. A real target child records its
identity, and a caller mutation back to the builder Python fails the
regression. Both downloaded SDK wheels match uv.lock and pass inventory
and availability checks. Ruff and added-comment checks passed.
No full native package or signing run is claimed.
Commit builds use their own immutable namespace and the shared signed
transport. Publish each receipt after its files, and fetch shared files
once only when their receipt records agree.
Summary links retain the full object key. A completed row requires its
own validated receipt and listed object. Missing, corrupt and ambiguous
results remain distinct. Include both universal Windows bundles.
Verified: 85 tests passed across transfer, rendering, candidate and
promotion paths. The subprocess test fetches the rendered download URL
from loopback HTTP. Ruff and added-comment checks passed.
No live R2 writes, workflow dispatch or native package build ran.
The commit-build CLI and workflow changes remain separate drafts.
The version writer only replaced an existing assignment. Modules
without the field therefore lost the Git-count anchor used for
package distance reporting.
Append the missing assignment and retain updates to an existing one.
A real temporary repository verifies that the emitted count equals
the resulting release commit, with other version files still aligned.
The revision and canary gates passed 20 tests. No Nix build or real
release operation was run.
Icon generation reported errors but returned success. Aggregate target
write and verification failures into exit 1, while processing later
targets. Keep rendering dependencies in the isolated build group.
A synchronous open error consumed a slot. Returning that slot alone
still let a deferred throw suppress an earlier callback and stall the
queue. Schedule queue draining outside completed callbacks and retain
the original exception.
Verified real clean-source generation and structural checks for all
35 targets, plus a real write failure through the Node/uv runner.
Native Windows child tests cover immediate and deferred open errors.
POSIX descriptor-limit cases are explicit skips here, not passes.
No signed package or native macOS acceptance is claimed.